Skip to main content

folkfox

Skip to main content
Skip to content
SEO AND GEO

Your search box is an open gate, and Googlebot walked through it

Nobody audits the search box. It sits in the header, quietly generating a URL for every phrase anyone has ever typed, including the phrases typed by people who do not wish you well.

Quick answerInternal site search pages can consume crawl budget at scale and, when spammed, may be flagged as hacked content in Search Console. Blocking them in robots.txt is the fix Google itself recommends.
SECTION 01

The open gate nobody put in the technical seo audit#

crawl budget

A fox tests a fence before it trusts it. Most sites never test the one gap that stays permanently open: the internal search box, which mints a fresh URL for every phrase a stranger cares to type.

The prompt for this is a conversation between Google's John Mueller and Martin Splitt, covered by Search Engine Journal on 31 July 2026. Mueller described the risk of letting people search for things wholly irrelevant to your website when your search results page then includes those terms. The consequence he named is sharper than a ranking wobble: Google might flag that as hacked, and you might see it in Search Console as something flagged as hacked.

Sit with that for a moment. Not demoted. Not filtered. Flagged as hacked, on a site nobody breached, because a spammer used your search box as a printing press and your server obligingly published the results. Spammers prowl for exactly this: a door left ajar in an otherwise tidy den.

Why crawl budget is the first thing to go#

Long before any trust flag appears, the crawl budget bleeds. Google's own guidance on managing crawling of faceted navigation URLs is unusually blunt about the mechanism: parameter-based implementations can generate infinite URL spaces, and crawlers cannot tell whether such URLs are useful without crawling them first, so they access a very large number before concluding the URLs are in fact useless.

That is the whole problem in one sentence. A crawler cannot know a page is worthless until it has spent the crawl budget finding out, and a search box offers an unlimited supply of pages to find out about. Googlebot follows every trail it is given, patiently, without ever asking whether the trail leads anywhere.

SECTION 02

What the evidence actually shows#

Google has said for years that low-value URLs cost you. In the 2017 post from Gary Illyes on what crawl budget means for Googlebot, Google states that according to its analysis, having many low-value-add URLs can negatively affect a site's crawling and indexing. Worth noting honestly: that analysis is asserted rather than published, so treat it as Google's stated position rather than as a study.

What makes that post unusually useful is the ranking. Google lists the categories in order of significance, and the order tells you where to point a technical seo audit first.

Google's own order of significance
Faceted navigation and session identifiers
1st
On-site duplicate content
2nd
Soft error pages
3rd
Hacked pages
4th
Infinite spaces and proxies
5th
Low quality and spam content
6th
Rank order only, as published by Google in 2017. The bar lengths encode position in Google's list, not measured impact, because Google published no magnitudes.

Google's closing line in that post is the one to quote at a sceptical engineering lead: wasting server resources on pages like these will drain crawl activity from pages that do actually have value, which may cause a significant delay in discovering great content on a site. Delay in discovery is the real cost. Your best new page does not rank badly, it simply waits.

For a measured figure rather than an asserted one, the best public record is the engineering case study from the Greek marketplace Skroutz, published on its engineering blog. Its team built Kibana-based log analysis covering more than 25 million pages, cross-referenced against traffic statistics and at least ten months of historical log data. The finding: more than 50% of daily crawl budget was spent on internal search pages, and most of those pages had no traffic at all.

Where one marketplace's crawl went
Donut chart showing more than half of daily crawl activity going to internal search pagesInternal search pages: 50%Everything else: 50%50%+
Internal search pages 50%Everything else 50%
Measured on a single site: Skroutz reported that more than half its daily crawl budget went to internal search pages with little or no traffic. One documented case study by the site's own engineering team, not controlled research.

After the cleanup, Skroutz reports its indexed URL count fell from roughly 25 million to 7.6 million. Read that as a shape rather than a target: this is one marketplace with one architecture, documented by the team that fixed it. It is still the most detailed public account of internal site search seo waste anyone has published.

Indexed URLs before and after
Indexed URLs before and afterSkroutz's reported indexed URL count fell from about 25 million to 7.6 million after blocking internal search and filter URLs. Single-site figures reported by its own engineers.Before cleanup: 25After cleanup: 7.62518.812.56.20Before cleanupAfter cleanup
Skroutz's reported indexed URL count fell from about 25 million to 7.6 million after blocking internal search and filter URLs. Single-site figures reported by its own engineers.
SECTION 03

When crawl waste becomes a trust problem#

Crawl budget is an efficiency argument, and efficiency arguments lose budget meetings. The trust argument wins them, and for regulated brands it is the one that matters.

Search Console's Security Issues report carries flags including hacked content injection, code injection and malware. Google describes hacked content as any content placed on your site without your permission because of security vulnerabilities, and says it tries to keep such content out of search results. Content injection is defined as a hacker adding spammy links or text to your pages.

Now look at that definition next to an unprotected search box. A third party causes spammy text to appear on a URL of your domain, without your permission, through a weakness in how your site handles input. From the outside, the fingerprint is identical. That is why Mueller's warning is not hyperbole.

There is a second exposure in Google's spam policies for web search, which list under doorway abuse the creation of substantially similar pages that are closer to search results than a clearly defined, browseable hierarchy. An indexable internal search space is precisely that, at scale, generated automatically.

That last bullet is the trap that catches careful teams. It is the reason a technical seo audit finds so many sites that believe they solved this years ago.

SECTION 04

Robots.txt or noindex, and why the answer is not both#

This is where most implementations go wrong, and the failure is silent, which is the worst kind.

Google's documentation on the noindex rule states that for it to be effective, the page must not be blocked by a robots.txt file and must otherwise be accessible to the crawler. Block the URL and the crawler never reads the noindex. The page can still surface in results, and you will believe you fixed it.

Which control you want depends on what you are protecting. If the goal is crawl budget, robots.txt wins, and Google says so directly in its crawl budget guidance for large sites: do not use noindex, because Google will still request the page and then drop it when it sees the rule, wasting crawling time. If the goal is removing an already-indexed page, you need it crawlable and carrying noindex, at least until it drops out.

Disallow plus noindex, applied together

The robots.txt rule stops the crawler reaching the page, so the noindex tag is never read. Already-indexed search URLs stay indexed indefinitely, and the team records the ticket as closed.

Noindex first, then disallow

Leave the URLs crawlable with a noindex header until Search Console shows them dropping out of the index. Only then add the robots.txt disallow to stop the crawl waste permanently.

For newly launched search spaces that were never indexed, skip straight to the disallow. Google's 2014 guidance on faceted navigation practices named this exact scenario as a worst practice: converting user-generated values into possibly infinite crawlable URLs. Its suggested remedy was to place user-generated values in a separate directory and disallow crawling of that directory. The post now carries a staleness banner, so pair it with the current documentation rather than citing it alone.

The separate-directory trick remains the cleanest engineering answer. One path prefix, one robots.txt line, no pattern-matching gymnastics, and a rule a new developer can understand in four seconds. It is the difference between fencing a whole thicket and fencing one gap in it.

Choosing between robots.txt and noindex by the outcome you actually want, with the failure each choice produces when applied to the wrong job.
Your goalCorrect controlWhat goes wrong otherwise
Stop crawl waste on new search URLsRobots.txt disallowNoindex still costs a crawl on every URL
Remove already-indexed search URLsNoindex, kept crawlableA disallow freezes them in the index
Both, on a legacy siteNoindex first, disallow after removalApplying both at once cancels the noindex
Consolidate near-duplicate filtersCanonical tagsDuplicate crawling continues indefinitely
Stop a live spam injectionDisallow now, noindex after cleanupWaiting for deindexing leaves spam visible

Canonicals deserve a mention because they are the tool people reach for by reflex. Google's guidance on consolidating duplicate URLs frames them as a way to avoid spending crawling time on duplicate pages, so Googlebot can spend time on new or updated pages instead. Useful for near-duplicate filter combinations. Useless against an unbounded search space, because a canonical still requires the crawl.

SECTION 05

Five fixes, in the order a fox would do them#

None of this is a replatform. All of it is an afternoon, and the first two moves recover most of the wasted crawl budget on their own.

Start by measuring, because a crawl waste argument without a log file is an opinion. Pull server logs, segment Googlebot requests by URL pattern, and count how many hit your search path. If the share resembles the Skroutz figure, you have your business case and you did not have to borrow anyone else's numbers to make it. The quarry here is not a ranking, it is the crawl your good pages are not getting.

Then close the gate, carefully#

Move search results to a dedicated path prefix if they are not already on one, disallow that prefix in robots.txt, and confirm in Search Console that crawl requests to it fall away. Where search URLs are already indexed, run the noindex step first and wait for removal before adding the disallow, for the reason above.

Then take the spam angle seriously, because faceted navigation seo hygiene and security hygiene are the same job here. Cap query length, reject queries containing markup or external URLs, and return a clean empty state rather than echoing the raw query into the page title. Echoing user input into a title tag is how a search page becomes a spam page with your brand on it.

What to watch after you ship

Crawl share on search paths

0%

Target zero within a fortnight of the disallow landing.

Indexed search URLs

0

Should fall steadily if noindex ran before the disallow.

Security issues open

0

The number that must stay at zero. Check monthly, not annually.

Keep an eye on the Manual Actions report as well as Security Issues. They are separate reports covering separate problems, and teams routinely check one and assume it covers both.

The wider lesson is one we keep meeting from different directions. Crawling is a budget, attention is a budget, and every automated surface you leave open spends both on your behalf. We made the same argument about AI crawlers in blocking AI crawlers closed one gate, and about measurement in search traffic fell 34 per cent.

For regulated brands there is a final reason to care. A trust flag on your domain is not a technical footnote, it is a conversation with a compliance officer, and it arrives without warning. The same reasoning drove our read of Google's review guidance in Google just banned the reviews Britain outlawed last year.

If you want the crawl budget work done rather than described, that is what folkfox SEO and GEO services does, and the same discipline runs through iGaming marketing and healthcare marketing, where a hacked flag is a licensing conversation rather than a ranking one.

Questions

Frequently asked questions#

Should internal site search pages be indexed by Google?

Generally no. They duplicate existing content, generate an unbounded number of URLs, and Google's spam policies treat pages closer to search results than a browseable hierarchy as doorway abuse. Block them from crawling once any indexed versions have been removed.

Can a search box really get my site flagged as hacked?

Google's John Mueller has said that where spammers exploit an indexable search box, Google might flag the site as hacked, and that this can appear in Search Console. No breach is needed, because the visible fingerprint matches injected spam content.

Should I use robots.txt or noindex on search pages?

Robots.txt to stop crawl waste, noindex to remove pages already in the index. Never both at once: a disallowed page is never crawled, so its noindex rule is never read, and the page can stay indexed indefinitely.

How much crawl budget do internal search pages waste?

It varies by site. The clearest public figure comes from the marketplace Skroutz, whose engineers reported that more than half their daily crawl budget went to internal search pages, most with no traffic. Measure your own logs rather than assuming that ratio.

Do small sites need to worry about crawl budget?

Crawl budget rarely limits small sites, but the trust and quality exposure applies at any size. A spammed search space can produce indexed junk and a security flag on a hundred-page site just as easily as on a hundred-thousand-page one.

What is the quickest safe fix?

Move search results onto a dedicated path prefix, add a noindex header, wait for the indexed URLs to drop out of Search Console, then disallow that prefix in robots.txt. One path, one rule, and no pattern-matching to maintain.

Googlebot still crawls pages before noindex exclusion, so the robots.txt and noindex are not interchangeable.

Keep reading

Read more on this topic#

Want to know what your crawl is actually spending on?

folkfox runs log-led technical audits for regulated brands: crawl waste, index bloat, and the trust flags that arrive without warning, reported in language your engineering lead will action.