Faceted navigation and SEO: avoiding crawl traps
· 6 min read
Here is a sum worth doing before you launch filters on a store. A clothing category with 10 colours, 8 sizes, 6 brands, 5 price bands and 4 sort orders can produce well over 100,000 distinct URLs once combinations and ordering are counted. Almost all of them show near-identical products. Search engine crawlers do not know that in advance, and they can spend their time on those pages instead of your real products. That, in short, is the problem faceted navigation SEO has to solve.
This article is about the search engine side of filters. It assumes the filters themselves already work well for shoppers.
How filter URLs multiply
Most stores record filter selections in the URL, for example /shoes?colour=black&size=9. That is good for shoppers, who can bookmark and share results. The trouble comes from the variations:
- Different orders of the same parameters: ?size=9&colour=black versus ?colour=black&size=9
- Multi-select values: ?colour=black,blue, ?colour=blue,black
- Sort and view options: ?sort=price_asc, ?view=grid
- Pagination layered on top of every combination
- Session IDs or tracking parameters appended by other tools
Each variation is a separate URL to a crawler. If filter links are plain crawlable links, Googlebot can discover them endlessly. This is what SEOs call a crawl trap.
Why crawl traps hurt a store
Google's documentation notes that crawl budget mostly matters for large sites or sites that generate many URLs, and faceted navigation is one of the classic causes. The practical effects are:
- New products and price changes are discovered and refreshed more slowly
- Near-duplicate pages appear in the index, sometimes ranking instead of the main category
- Server load rises as bots request thousands of filter pages
In Google Search Console, the Pages report often shows the symptoms: large numbers of URLs under "Duplicate without user-selected canonical" or "Crawled – currently not indexed", many of them with filter parameters. The Crawl stats report under Settings can also show crawlers spending heavily on parameter URLs.
Deciding which filter pages deserve indexing
Not every filter page is waste. Some combinations match what people actually search for. "Black running shoes" or "cotton sarees" are real queries, and a filtered page targeting them can rank well if it is treated like a proper landing page.
A simple test for each candidate
- Is there meaningful search demand for this attribute plus category? Check with keyword tools or Search Console queries.
- Does the filtered page contain enough products to be useful, not just one or two?
- Would the page be substantially different from the parent category?
- Can you give it a unique title, heading and a short intro?
If the answer to all four is yes, make it an indexable page, ideally at a clean static URL such as /shoes/black-running-shoes. Typically this means a limited set of single-facet pages, such as one colour or one brand. Multi-facet combinations, price ranges, sort orders and view modes almost never qualify.
Tools for controlling filter URLs
There is no single switch. Each tool does a different job, and mixing them up causes problems.
Canonical tags
A rel="canonical" tag on a filtered page pointing to the main category tells Google which version you prefer. It is a strong hint, not a command. It helps consolidate signals from near-duplicates, but crawlers may still fetch those URLs, so canonicals alone rarely fix crawl waste. Only canonicalise a page to one that shows substantially the same content.
Meta robots noindex
A noindex tag keeps a page out of search results. Google must crawl the page to see the tag, so noindex does not save crawl budget directly. It suits filter pages you want reachable but not listed.
Robots.txt disallow
Blocking parameter patterns in robots.txt, for example lines disallowing URLs containing sort= or multiple parameters, stops compliant crawlers from fetching them at all. This is the most direct way to save crawl budget. Two cautions: blocked URLs can still be indexed without content if other sites link to them, and Google cannot see a noindex or canonical tag on a page it is not allowed to crawl. Do not combine a disallow with a noindex you expect Google to read.
Not creating crawlable links
For low-value options such as sort order, the cleanest approach is not to expose them as standard links at all. Applying them through form controls, or through JavaScript that updates results without generating crawlable href links, reduces discovery at the source. Make sure shoppers can still use and share the results.
Google Search Console's old URL Parameters tool was retired in 2022, so parameter handling now has to be done on the site itself.
A workable setup for most small stores
- Keep one consistent parameter order so the same selection always produces the same URL
- Give a small, chosen set of high-demand facet pages clean URLs, unique titles and intros, and let them be indexed
- Canonicalise other single-filter URLs to the parent category
- Block sort, view and multi-filter parameter patterns in robots.txt once you are sure nothing valuable uses them
- List only the category and chosen facet pages in your XML sitemap
- Strip session and tracking parameters from internal links
Test robots.txt changes carefully before deploying. A pattern that is too broad can block real category or product pages. After changes, watch the Pages and Crawl stats reports over the following weeks.
Checking your store before crawlers find the problem
Open a category, apply three filters and a sort, and look at the URL. Then view the page source for the canonical and robots tags. If every combination is crawlable, indexable and self-canonical, you have a trap waiting to happen.
Fixing this usually involves templates and server configuration, so it is a job for whoever maintains the store code. If that is unclear, our team handles technical SEO like this during e-commerce development, and you can ask a specific question through our support page.
Frequently asked questions
Should I noindex all filtered category pages?
Not all of them. Filter pages that match real search demand can be valuable landing pages. Noindex or block the combinations with no search value, and give the useful ones proper content.
Can I use canonical and noindex on the same page?
Google advises against sending mixed signals. A canonical says "treat this as a copy of another page", while noindex says "drop this page". Choose the one that matches your intent.
Does AJAX filtering solve faceted navigation SEO issues?
It can reduce crawlable URLs, but only if filter states are not exposed as ordinary links. Check that Google can still reach every product through category pages and pagination.
How long does it take for Google to drop old filter URLs?
It varies with site size and crawl rate. For many stores it takes weeks to a few months for the Pages report numbers to settle after changes.
Thinking about a website?
See what a package covers and what it costs, or ask us about your own project.