Skip to main content

folkfox

Skip to main content
Skip to content
SEO AND GEO

Blocking AI crawlers closed one gate. Googlebot was behind it

One toggle in a CDN dashboard, one 403 on a sitemap, and a site that ranked on Monday is thinning by Friday. Nobody on the marketing side gets an email about it.

Quick answerBlocking AI crawlers can also block search. Googlebot and bingbot crawl for indexing and AI at once, so a training block can return 403s, and Google removes 403 URLs from its index.
Section 01

The 403 that nobody ordered#

blocking ai crawlers

A fox does not torch the whole hedgerow to be rid of one rabbit. It watches, waits, works out which gap the quarry actually uses, and closes that single gap. Blocking AI crawlers has become the marketing equivalent of burning the hedgerow flat, and the creature you never meant to catch is usually the one paying the bills.

On 4 August 2026, Search Engine Journal reported a site owner posting in the r/TechSEO community whose sitemap began returning HTTP 403 to Googlebot and bingbot once Cloudflare's AI Training block was set to block. The owner's account was blunt: with the setting enabled the sitemap fetch returns a 403, and disabling it makes the 403 disappear. Google's John Mueller asked the poster to send him a message so he could investigate the case directly.

Be fair to the vendor, because this is a hard problem and publishers begged for the control. Cloudflare shipped a granular version of it. Its AI Crawl Control documentation describes setting allow or block rules for individual crawlers, monitoring which AI services arrive, tracking robots.txt compliance, and even charging for access through Pay Per Crawl. The criticism here is not of the feature. It is of the blunt use of a sharp tool, by people who cannot see what the swing costs.

Bot policing tightens, and blocking AI crawlers is one front of it#

The mood is not confined to CDNs. On the same day, Search Engine Land reported a limited, unconfirmed test in which Google swapped its usual bot check for a prompt asking searchers to sign in to see more results, spotted and shared by Kamlesh Shukla. Google has not confirmed it. Read that as weather rather than forecast: everyone is narrowing the definition of a legitimate visitor, and blocking AI crawlers is the version of that instinct which lands on your own infrastructure, on your own budget.

Section 02

One crawler, two appetites: why blocking AI crawlers catches search too#

Here is the mechanism, and it is the whole argument in a sentence: the crawler that trains the model and the crawler that fills the index are increasingly the same animal in one coat. That is precisely why blocking AI crawlers is not a surgical strike. Where a bot is single-purpose, the block is clean and correct. Where a bot is multi-purpose, the block becomes collateral, and the collateral is your organic revenue.

Google actually offers the clean cut. Its list of common crawlers gives Google-Extended its own robots.txt token, and states plainly that Google-Extended governs whether crawled content may be used to train Gemini models and to ground responses in Gemini Apps and Vertex AI, and that it does not impact a site's inclusion in Google Search nor act as a ranking signal. Disallow that one token and you have declined the training use while keeping the index. Surgical, free, and available today in a text file.

OpenAI splits the same three ways. Its crawler documentation separates GPTBot, which feeds foundation models, from OAI-SearchBot, which surfaces sites inside ChatGPT search features, from ChatGPT-User, which fetches a page because a person asked it to. OpenAI states each setting is independent of the others. Anthropic draws the same lines in its crawler guidance: ClaudeBot for training, Claude-SearchBot for search quality, Claude-User for live user requests.

So blocking AI crawlers with a scalpel is entirely possible whenever the vendor hands you a scalpel. The trouble starts with the bots that refuse to be split into separate skins.

Cloudflare's announcement of new AI traffic options, published on 1 July 2026, sorts crawling into three purposes: Search, which collects or indexes content so it can answer questions about it later; Agent, which acts in real time on a person's behalf; and Training, which takes content to train or fine-tune a model. Where one crawler serves more than one purpose, the most restrictive rule wins.

One switch, every crawler barred

A zone-wide edge rule handles blocking AI crawlers everywhere at once. GPTBot and ClaudeBot stop, and so does every multi-purpose crawler carrying a search job in the same coat. The sitemap answers 403. Nothing in the marketing stack says a single word about it, and the first symptom surfaces weeks later as a soft, unexplained slide in impressions.

Named tokens, kept receipts

robots.txt disallows Google-Extended, Applebot-Extended, GPTBot and ClaudeBot by name. Googlebot, bingbot and OAI-SearchBot stay explicitly allowed. Edge rules challenge unverified traffic rather than blanket-blocking declared search agents, and every change carries a date, an owner and a rollback note. Same intention, a tenth of the risk, and a trail you can follow backwards.

Crawl requests per referral visit
Anthropic
73,000:1
OpenAI
1,700:1
Google
about 14:1
Publishers are not paranoid, they are counting. Ratios as Cloudflare measured them in June 2025 and published in its July 2025 post on managed robots.txt. Bars are scaled by order of magnitude because the range spans four of them; the figure printed on each bar is the published ratio.

Look at that spread and the switch-flipping instinct makes perfect sense. Tens of thousands of fetches for a single visit is not a fair forage, it is a warren being stripped. The instinct is sound and the implementation is where sites get hurt, because blocking AI crawlers at the zone level is a hammer swung in a dark den, aimed at a shape you cannot quite see.

Section 03

The crawler table: who is who, and what a block really costs#

Every serious conversation about blocking AI crawlers should start with a table rather than a toggle. Print this one, take it to whoever owns the CDN, and make them say out loud which rows they intend to close. Half the risk evaporates the moment the decision is named crawler by crawler instead of category by category.

Every crawler token named in vendor documentation, what each one is for, and what blocking AI crawlers row by row actually costs a marketing team.
Crawler tokenWhat it is forWhat blocking it costs you
GooglebotGoogle Search indexing. Listed as a common crawler that always respects robots.txt for automatic crawls.Your presence in Google Search. Never block this one.
Google-ExtendedControls whether crawled content may train Gemini models and ground responses in Gemini Apps and Vertex AI.Grounding and training use. Google states it is not a ranking signal and does not impact Search inclusion.
Googlebot-NewsNews crawling. It has no separate user agent string and reuses Googlebot strings.Google News surfaces. Rarely worth blocking.
GoogleOtherA separate common-crawler token with its own robots.txt entry, distinct from Googlebot.Varies by product. Read the entry before you close it.
bingbotMicrosoft's search crawler, verified by reverse DNS to a search.msn.com hostname.Your presence in Bing, and in surfaces that lean on Bing's index.
GPTBotOpenAI's crawler for making generative foundation models more useful and safe.Training use of your content. No ChatGPT search penalty on its own.
OAI-SearchBotSurfaces websites in ChatGPT's search features.Visibility inside ChatGPT search. Block only if you truly want out.
ChatGPT-UserFetches a page when a person asks ChatGPT about it, or via GPT Actions.Live user-initiated fetches. Blocking makes you unquotable on demand.
ClaudeBotAnthropic's crawler collecting content that may contribute to training.Training use only. Anthropic documents it separately from search.
Claude-SearchBotNavigates the web to improve search result quality.Indexing for Claude's search answers. A quiet GEO cost.
Applebot-ExtendedNamed alongside Google-Extended in Cloudflare's managed robots.txt directives.Apple's AI training use, per Cloudflare's own published wording.

Never trust the string on its own, because anyone can forge a user agent header. Bing's webmaster team tells site owners to run a reverse DNS lookup on the requesting address, confirm the hostname ends in search.msn.com, then run a forward lookup back to the same address. Bing deliberately publishes no crawl IP ranges, because they change and hardcoded lists end up blocking bingbot by accident.

Google names the same three signals in its crawler overview: the user agent header, the source address, and the reverse DNS hostname. That page also splits Google's fleet into common crawlers, special-case crawlers and user-triggered fetchers, and notes that AdsBot can ignore the global robots.txt rule with publisher permission. Sensible ai bot management seo work starts by knowing which of those three buckets a request sits in.

The robots.txt reality, mid-2025
Bar chart comparing robots.txt adoption among top domains, GPTBot reach across websites, and the share of robots.txt files disallowing GPTBotTop 10k with robots.txt: 37Sites GPTBot reached: 29Files blocking GPTBot: 7.837%27.8%18.5%9.2%0%Top 10k with rSites GPTBot rFiles blocking
Most sites never wrote the file that would have made blocking AI crawlers surgical in the first place. Figures published by Cloudflare in July 2025 from measurements across its own network; the three bars use different populations, so read each label.
Section 04

The failure is silent, and silence is the expensive part#

A 403 is not a polite refusal, it is a deletion order with better manners, and it is how blocking AI crawlers turns into losing pages. Google's documentation on HTTP and network errors states that all 4xx errors except 429 are treated the same way: the crawler tells the next processing system that the content does not exist, and for Search the indexing pipeline removes the URL from the index if it was previously indexed. Read that sentence to whoever owns the edge configuration.

Search Console will tell you eventually, in a report most teams open far too rarely. The page indexing report carries a dedicated status for pages blocked by a forbidden response, and Google's guidance is direct: Googlebot never provides credentials, so a 403 aimed at it means the server is returning that error incorrectly. The remedy Google gives is to admit visitors who are not signed in, or to allow verified Googlebot requests through without a credential challenge. Blocking AI crawlers should never produce that status for a search bot.

How the silence unfolds
Six steps, no alert. Blocking AI crawlers fails quietly, and the gap between the toggle and the question is where the money goes.

Two comfortable myths worth a clean kill#

First, robots.txt is not an index control. Google's introduction to robots.txt says outright that it is not a mechanism for keeping a page out of Google, and a disallowed URL that other sites link to can still appear in results, simply without a description. Second, robots.txt is advisory. The standard itself, RFC 9309, published in September 2022, binds only the crawlers that choose to be bound.

There is a genuinely nasty detail buried in that RFC which almost nobody has read. If robots.txt returns a 5xx, a compliant crawler must assume complete disallow. If it returns a 4xx, the crawler may access any resource on the server. So a blanket edge rule that answers 403 to any unrecognised agent does not tighten your rules, it deletes them for every bot still willing to knock. Blocking AI crawlers badly can leave you looser than blocking nothing at all.

The shape of a silent block
The shape of a silent blockLine chart showing indexed URLs falling before impressions across seven weeks after a silent crawler block1007550250W1W2W3W4W5W6W7Indexed URLs: 100Indexed URLs: 100Indexed URLs: 96Indexed URLs: 84Indexed URLs: 67Indexed URLs: 51Indexed URLs: 44Impressions: 100Impressions: 99Impressions: 97Impressions: 90Impressions: 76Impressions: 60Impressions: 49
Indexed URLsImpressions
Illustrative, not measured: these lines are a modelled shape rather than data from any site, drawn to show how indexed URLs fall first and impressions follow with a lag, which is exactly why the cause is hard to trace back to a dashboard toggle.

Crawl capacity makes the lag worse, though not quite in the way most people picture it. Google's guidance on managing crawl budget for large sites explains that server errors, rate-limiting signals and slower responses pull the crawl capacity limit down, and that crawl demand differs per crawler while the capacity itself is shared.

A forbidden response is a different beast in that respect, because Google's error documentation states that 4xx codes other than 429 have no effect on crawl rate. So the block does not slow the crawl politely, it strips the pages out while the crawler keeps knocking. Documentation on crawl budget is worth a careful read before anyone starts throttling bots at the edge, because a site that crawls slowly recovers slowly.

A block is a business decision wearing a network setting's clothes. Someone in marketing should be in the room when it is made.
folkfox, on ai crawler blocking and who signs it off
Section 05

A five-step audit to run before 15 September 2026#

None of this needs a replatform, a retainer or a heroic sprint. It needs one afternoon, one shared document, and one person willing to ask an awkward question in a channel where nobody usually asks. Run the audit now, while blocking AI crawlers is a choice you are making rather than a default you inherited on a date somebody else set.

The five-step crawler audit
Fetch as the bots

Request your sitemap, homepage and three money pages as Googlebot and as bingbot from outside your network. Anything other than a 200 is the whole story, and a googlebot 403 error on a sitemap is a five-alarm finding, not a curiosity.

Inventory the rules

List every live bot rule on the zone, its owner, its date and its intention. Rules with no named owner get closed or claimed. Undocumented edge rules are the undergrowth this problem hides in.

Name the tokens

Rewrite blanket blocks as named robots.txt disallows: Google-Extended, Applebot-Extended, GPTBot, ClaudeBot. Explicitly allow Googlebot, bingbot and OAI-SearchBot so nobody has to guess your intention later.

Check the defaults

Open your CDN security settings and confirm what happens on 15 September 2026. Decide the opt-out deliberately, record who decided, and diarise a review for the week after the date lands.

Wire an alarm

Alert on non-200 responses to verified search-engine user agents, and put the page indexing report on a monthly calendar invitation with a named owner. Blocking AI crawlers without monitoring is a snare set in your own path.

The reporting layer matters as much as the rules. If impressions slide and nobody can say whether the cause was an algorithm, a competitor or a checkbox, the argument in the review meeting will be settled by whoever speaks with most conviction rather than most evidence. That is how good sites lose quarters. Our work on AI visibility and citation reporting exists for exactly that reason.

There is a strategic reading here too, and it is not the doom-laden one. Being visible to AI systems is now part of the same discipline as being visible in search, which is why we treat them as one brief in SEO and GEO services rather than two competing workstreams. Blocking AI crawlers wholesale opts you out of the answer layer as surely as a noindex tag opts you out of the results page, and the pieces on SEO versus GEO and the hidden costs of opting out make the same point from the other direction.

Regulated categories carry the extra weight, as ever. An operator in iGaming marketing or FinTech marketing that vanishes from the index does not merely lose traffic, it loses the ability to correct what a model says about its licence while somebody else's page fills the gap. Sound content marketing assumes the content can be reached. Check the gate before you trust the trail.

The fox lesson, then, and it costs nothing to learn twice. Do not close the whole hedgerow. Find the gap the quarry actually uses, close that one, mark it on a map somebody else can read, and come back in September to check the gate still swings the way you left it.

Aakash Gupta
"On September 15, Cloudflare will start cutting off Google traffic for millions of its smallest customers by default. And those customers won't leave. They've been begging for this...
2026View on X
Questions

Frequently asked questions#

Does blocking AI crawlers stop Google from indexing my site?

It can. Google-Extended is a separate token, and disallowing it does not affect Search. A blanket edge rule that returns 403 to unrecognised agents is different: Google treats a 403 as content that does not exist and removes previously indexed URLs, so the damage appears as pages quietly leaving the index rather than as an error anyone gets alerted to.

What actually changes on 15 September 2026?

Cloudflare has said its defaults will block Training and Agent crawlers on pages carrying ads for new domains, while Search stays allowed. Because multi-purpose crawlers such as Googlebot, Applebot and BingBot are judged by every purpose they serve, customers who block Training block those bots too. An opt-out sits in security settings before the date.

How do I block AI training without losing search traffic?

Name the tokens instead of the category. Disallow Google-Extended, Applebot-Extended, GPTBot and ClaudeBot in robots.txt, and explicitly allow Googlebot, bingbot and OAI-SearchBot. That is blocking AI crawlers with a scalpel rather than a hammer, and it keeps every indexing crawler you actually depend on.

Is blocking AI crawlers ever the right call?

Often, yes. If your content is your product, blocking AI crawlers that only train models is a defensible commercial decision, and vendors now publish separate tokens precisely so you can make it. The failure mode is not the intention, it is the blunt instrument: category-level blocks that sweep up multi-purpose search crawlers you never meant to touch.

How do I check a crawler really is Googlebot or bingbot?

Verify the source rather than the string, because user agents are trivially forged. Run a reverse DNS lookup on the requesting address, confirm the hostname belongs to the search engine, then run a forward lookup back to the same address. Bing points at hostnames ending in search.msn.com; Google lists the user agent, the address and the reverse DNS hostname as its three signals.

Will robots.txt keep my pages out of AI answers?

Only for crawlers that honour it. The robots exclusion standard is voluntary, and Google states that robots.txt is not a mechanism for keeping a page out of Search at all, since a disallowed URL linked from elsewhere can still appear without a description. For genuine exclusion you need noindex, authentication or a rule enforced at the edge.

Who should own the decision about blocking AI crawlers?

Marketing, SEO and engineering jointly, with one named owner and a written record. The setting lives in infrastructure but the consequence lands on revenue, so a change should carry a date, an intention and a rollback note. Review it on a schedule rather than the first time somebody notices impressions sliding.

Keep reading

Read more on this topic#

Want the gate checked before September?

folkfox audits crawler access, edge rules and AI visibility together, so blocking AI crawlers stays a decision you made rather than a default you inherited, with reporting that flags the moment a setting starts costing you the index.