An online shop with about 4,200 products had 31,900 pages listed in Google by August. Roughly 27,000 of them didn't exist.
In May, Google's count of its pages stood at 5,500, which is the catalogue plus a few dozen category and information pages. By August it was 31,900. Nobody had added a page.
Where the pages came from
Automated bots had been hammering the site with made-up addresses. Take a real product page, tack a search term or a sort order onto the end, and the shop platform obligingly served a full page for it. Every one of those addresses counts as a new page, and Google, doing its job, filed each one. Six months of that and the index was five parts junk to one part shop.
Google only spends so much time on any one site. Over 90 days it made 609,000 requests to this one, half of them for addresses it had never seen before, which on a catalogue that barely changes means junk. And the server built every one of those pages for a visitor who was never a person.
The first fix, and why it backfired
The obvious answer is to tell Google to keep out. Every website can carry a small file called robots.txt, a list of the addresses search engines are asked not to visit. I added the junk patterns to it in July, and for five days it looked like the right call.
Then a new line appeared in Google's report: "Indexed, though blocked by robots.txt". It started at 158 and kept climbing. The reason sits in Google's own documentation, in so many words: robots.txt is not a way of keeping a page out of Google. It stops Google reading a page. It doesn't stop Google listing it.
Worse, every junk address already carried a small note in its code pointing at the real product page, and Google had been using that note to tidy up. Now it couldn't read the note. The junk pages already in the index were frozen there: listed, unreadable, and impossible to remove.
What worked
The fix was the opposite. Take the sign down, let Google in, and give it something it can act on when it arrives. A junk address on a product page now sends Google straight to the real one, a permanent redirect. The shop's own search results pages, which turned out to be the only ones being listed in bulk, carry an instruction not to list them. Both only work if Google visits the page, which is exactly what the block had been preventing.
Ten days later Google had found 26,790 of the redirects. By 4 September, five weeks after the change, the page count was 5,220. Back where it was in May.
The block isn't gone for good. Once the index has stayed clean for a second reading, the keep-out rules go back in to save Google the wasted visits, because by then there's nothing left to freeze. Clean the index first, then cap the crawling.
The number that looked like bad news
Along the way, the report's "not indexed" figure climbed from 244,000 to 300,000, and the count of indexed pages fell by 27,000. Both look like a site in trouble. Both were the clean-up working, because a redirect or a "don't list this" instruction gets filed under "not indexed". Had I not known what correct looked like, 5,500 in May against 4,600 real pages, it would have been easy to read the fall as damage and undo the fix.
The figure that mattered didn't move. The shop appeared in Google searches between 4,500 and 5,000 times a day throughout: 756,000 appearances and 43,400 visits over six months, with no dip while 27,000 pages left the index. Whatever those pages were doing in Google's list, they weren't bringing anyone to the shop.
The rest of the once-over
The same visit turned up the kind of thing a site collects when nobody is looking. The home page was 6.9 MB to download; it's now 1.0 MB with the same pictures. A small icon that browsers ask for automatically had been requested 142,000 times and didn't exist. An old contact-page address had been visited 18,000 times and led nowhere. And the sitemap, Google's contents page for the site, had 914 entries with bad dates and one fault that broke the whole file.
None of it was dramatic. All of it was cheap to put right, and all of it had been costing the shop for a long time.
If your own report has a number you can't explain
Google's free report on your site is called Search Console, and most owners have never opened it. If something in there looks alarming, the first question is what correct looks like. For a shop it's roughly the catalogue. If the count is a long way above that, you have a junk problem, and the way out is to let Google in and tell it what to do with each address. Blocking feels safer. It's also the one move that guarantees the mess stays where it is.