
Quick answer
Index bloat is when Google indexes many more URLs than you want found, such as thin archives, filter combinations and parameter pages. Find it by comparing indexed pages in Search Console with your intended set. Then decide each URL’s fate: improve it, noindex it, canonicalise it, redirect it or remove it. It is not a penalty.
Ahrefs (2023) found that 96.55 per cent of the roughly 14 billion pages in its index get zero organic traffic from Google, a reminder that indexed pages are not the same as useful pages.
Index bloat happens when Google has indexed far more of your URLs than you actually want people to find, usually thin archives, filter combinations and parameter pages. This guide shows you how to spot it in Search Console, how to decide what each unwanted URL needs, and how to remove it without damaging the pages that earn traffic.
What index bloat actually is
Every site has a set of pages that deserve to appear in search: your services, products, categories and useful articles. Index bloat is everything else that Google has picked up on top of that set. It is not a penalty and there is no official threshold. It is simply a gap between the pages you intend to be indexed and the pages that are.
A small gap is normal. A large one, where low-value URLs outnumber your real content, is a signal that your site is generating pages faster than you are deciding what to do with them. That is a structural problem, not a content problem, and adding more articles will not fix it.
Why it hurts: crawling and quality signals
There are two separate costs, and it helps to keep them apart.
Wasted crawling
Google does not crawl every URL on every site constantly. On a large site, time spent fetching thousands of filter combinations is time not spent discovering new products or refreshing updated pages. For small sites this is rarely the main issue, but on e-commerce and large publishing sites it can slow how quickly important changes are picked up. The guide to crawl budget in SEO explains when this matters and when it does not.
Diluted quality signals
Google has said that its systems look at site-wide signals as well as individual pages. If a large share of your indexed URLs are empty tag pages, near-duplicate product variants or internal search results, the overall impression of the site is weaker than its best content deserves. There is also a simpler effect: when several thin URLs target the same topic, they compete with each other and none of them ranks well.
How to find index bloat
You are trying to answer one question: which indexed URLs would I not choose to show a searcher? Use these sources together, because none gives the complete picture alone.
Step 1: Count your intended pages
Write down roughly how many pages you want indexed. Count your published posts, pages, products and the categories you deliberately promote. This is your baseline. Without it, an indexed page count means nothing.
Step 2: Read the page indexing report
In Search Console, the page indexing report shows how many pages are indexed and lists the reasons other pages are not. Compare the indexed total with your baseline. If indexed pages are well above your intended count, open the indexed list and look at the URL patterns. Also look at the “not indexed” reasons: large numbers of “crawled, currently not indexed” or “duplicate without user-selected canonical” often point to the same sources that cause bloat.
Step 3: Run targeted site: searches
The site: operator gives rough, unreliable counts, so do not treat the number as accurate. It is useful for spotting patterns. Try searches such as:
site:yourdomain.com inurl:tagfor tag archivessite:yourdomain.com inurl:?for parameter URLssite:yourdomain.com inurl:pagefor deep paginationsite:yourdomain.com inurl:searchfor internal search results
Step 4: Crawl and compare
Crawl your site with a desktop crawler and export every indexable URL. Group them by folder or URL pattern. The groups that are large but contain nothing a searcher would want are your bloat sources.
The usual sources on WordPress and e-commerce sites
- Tag archives: tags used once or twice create pages that list a single post.
- Author and date archives: on single-author blogs these duplicate the main blog feed.
- Image attachment pages: a WordPress default that produces a page per uploaded image. The WordPress technical SEO checklist covers how to switch these off.
- Faceted filters: colour, size, price and brand filters that each generate a crawlable URL, multiplied by every combination.
- Sorting and tracking parameters:
?sort=price,?ref=and campaign tags that create copies of the same page. - Internal search results: if your search result pages are linkable and indexable, any query can become a page.
- Page builder templates: saved headers, popups and template library items that become public URLs. This is common on Elementor sites, which I cover in Elementor SEO technical fixes.
- Old migrations: previous URL structures still reachable because redirects were never set up. The website migration SEO checklist shows how to map and redirect old URLs properly.
Deciding what each URL needs
This is where most clean-ups go wrong. People pick one tool, usually noindex, and apply it to everything. Each option sends a different instruction, and the right one depends on whether the page has value, has a better equivalent, or should not exist.
| Situation | Best action | Why |
|---|---|---|
| Page is useful to visitors on the site but not in search (e.g. internal search results, thin tag archive) | noindex, keep it crawlable |
Visitors can still use it; Google drops it from the index once it recrawls and sees the tag |
| Page is a near-copy of another page (sort orders, tracking parameters, print versions) | Canonical to the main version | Consolidates signals onto the preferred URL while the copy keeps working |
| Page has an obvious replacement and has links or traffic | 301 redirect | Passes visitors and signals to the replacement |
| Several weak pages cover the same topic | Consolidate into one strong page, redirect the rest | One complete page usually serves searchers better than three partial ones |
| Page has no value, no links, no traffic and no replacement | Delete and return 404 or 410 | Removes it cleanly; Google drops it over time |
| Filter combinations nobody searches for | Stop generating crawlable links to them, then noindex or canonical what already exists | Fixes the source, not just the symptom |
Two rules I follow. First, use one instruction per URL. A page that is noindexed and also canonicalised elsewhere sends conflicting messages. Second, do not block a URL in robots.txt if you want Google to see a noindex tag on it. If Google cannot crawl the page, it cannot read the tag, and the URL can stay indexed. My explainers on canonical tags and robots.txt go deeper into how each one behaves.
A worked example
Illustrative example: imagine a clothing shop with around 400 products and 30 categories. Its owner expects roughly 450 indexed pages. Search Console shows several thousand. Looking at the indexed URLs, most contain ?colour=, ?size= or ?orderby=.
The fix happens in layers:
- Sorting parameters get a canonical pointing to the unsorted category, because they show the same products.
- Filter links that create endless combinations are changed so they do not generate new crawlable URLs for every click. A small number of filters that match real searches, such as a single colour within a category, could be turned into proper landing pages with their own content.
- Remaining filter URLs get a noindex tag and stay crawlable until Google has processed them.
- The sitemap is trimmed to products and real categories only.
- Progress is checked in the page indexing report over the following weeks.
Notice that nothing here involves deleting products. The goal is to stop the site manufacturing duplicates, not to shrink the shop.
Keeping it from coming back
Index bloat is almost always caused by a system: a plugin, a theme, a filter module or a habit. Fix the system or the problem returns.
- Agree a tagging policy and merge or remove tags that are used rarely.
- Check new plugins for any post types or archives they create.
- Review the page indexing report regularly and investigate any sudden rise in indexed pages.
- Keep your sitemap limited to canonical, indexable URLs, and remove any section that lists archives or parameters you have decided to exclude.
When the bloat is tangled up with JavaScript filters, migrations and conflicting plugins, it becomes a proper diagnostic job. That is a core part of a technical SEO audit and fixes engagement.
Questions people ask about index bloat
How long does it take for noindexed pages to drop out of Google?
It depends on how often Google recrawls those URLs. Popular pages may be processed quickly; rarely visited ones can take much longer. You cannot force a fixed timeline, but keeping the pages crawlable and linked from a temporary sitemap can help Google find the change.
Should I use the removals tool to fix index bloat?
No. The removals tool in Search Console hides URLs temporarily. It does not change how Google treats them long term. Use it only for urgent cases such as leaked private pages, and always pair it with a permanent fix.
Is a large number of “crawled, currently not indexed” pages a problem?
Not on its own. Google is telling you it saw those pages and chose not to index them. It becomes worth investigating when important pages are in that list, or when the list is dominated by a URL pattern your site should not be generating at all.
Does deleting thin pages always help rankings?
No, and you should not expect a guaranteed result. Removing low-value pages makes the site cleaner and easier to crawl, but pages with traffic, links or real purpose should be improved or merged rather than deleted.
If your indexed page count does not match what you think your site contains, it is worth finding out why before you publish more content. Ask me for a free audit and I will show you where the extra URLs come from and which fix suits each group.
