Blocking AI crawlers closed one gate. Googlebot was behind it
One toggle in a CDN dashboard, one 403 on a sitemap, and a site that ranked on Monday is thinning by Friday. Nobody on the marketing side gets an email about it.
By Katie Delaney · 2026-08-06 · 17 min read
The 403 that nobody ordered#

A fox does not torch the whole hedgerow to be rid of one rabbit. It watches, waits, works out which gap the quarry actually uses, and closes that single gap. Blocking AI crawlers has become the marketing equivalent of burning the hedgerow flat, and the creature you never meant to catch is usually the one paying the bills.
On 4 August 2026, Search Engine Journal reported a site owner posting in the r/TechSEO community whose sitemap began returning HTTP 403 to Googlebot and bingbot once Cloudflare's AI Training block was set to block. The owner's account was blunt: with the setting enabled the sitemap fetch returns a 403, and disabling it makes the 403 disappear. Google's John Mueller asked the poster to send him a message so he could investigate the case directly.
Be fair to the vendor, because this is a hard problem and publishers begged for the control. Cloudflare shipped a granular version of it. Its AI Crawl Control documentation describes setting allow or block rules for individual crawlers, monitoring which AI services arrive, tracking robots.txt compliance, and even charging for access through Pay Per Crawl. The criticism here is not of the feature. It is of the blunt use of a sharp tool, by people who cannot see what the swing costs.
Bot policing tightens, and blocking AI crawlers is one front of it#
The mood is not confined to CDNs. On the same day, Search Engine Land reported a limited, unconfirmed test in which Google swapped its usual bot check for a prompt asking searchers to sign in to see more results, spotted and shared by Kamlesh Shukla. Google has not confirmed it. Read that as weather rather than forecast: everyone is narrowing the definition of a legitimate visitor, and blocking AI crawlers is the version of that instinct which lands on your own infrastructure, on your own budget.
One crawler, two appetites: why blocking AI crawlers catches search too#
Here is the mechanism, and it is the whole argument in a sentence: the crawler that trains the model and the crawler that fills the index are increasingly the same animal in one coat. That is precisely why blocking AI crawlers is not a surgical strike. Where a bot is single-purpose, the block is clean and correct. Where a bot is multi-purpose, the block becomes collateral, and the collateral is your organic revenue.
Google actually offers the clean cut. Its list of common crawlers gives Google-Extended its own robots.txt token, and states plainly that Google-Extended governs whether crawled content may be used to train Gemini models and to ground responses in Gemini Apps and Vertex AI, and that it does not impact a site's inclusion in Google Search nor act as a ranking signal. Disallow that one token and you have declined the training use while keeping the index. Surgical, free, and available today in a text file.
OpenAI splits the same three ways. Its crawler documentation separates GPTBot, which feeds foundation models, from OAI-SearchBot, which surfaces sites inside ChatGPT search features, from ChatGPT-User, which fetches a page because a person asked it to. OpenAI states each setting is independent of the others. Anthropic draws the same lines in its crawler guidance: ClaudeBot for training, Claude-SearchBot for search quality, Claude-User for live user requests.
So blocking AI crawlers with a scalpel is entirely possible whenever the vendor hands you a scalpel. The trouble starts with the bots that refuse to be split into separate skins.
Cloudflare's announcement of new AI traffic options, published on 1 July 2026, sorts crawling into three purposes: Search, which collects or indexes content so it can answer questions about it later; Agent, which acts in real time on a person's behalf; and Training, which takes content to train or fine-tune a model. Where one crawler serves more than one purpose, the most restrictive rule wins.
One switch, every crawler barred
A zone-wide edge rule handles blocking AI crawlers everywhere at once. GPTBot and ClaudeBot stop, and so does every multi-purpose crawler carrying a search job in the same coat. The sitemap answers 403. Nothing in the marketing stack says a single word about it, and the first symptom surfaces weeks later as a soft, unexplained slide in impressions.
Named tokens, kept receipts
robots.txt disallows Google-Extended, Applebot-Extended, GPTBot and ClaudeBot by name. Googlebot, bingbot and OAI-SearchBot stay explicitly allowed. Edge rules challenge unverified traffic rather than blanket-blocking declared search agents, and every change carries a date, an owner and a rollback note. Same intention, a tenth of the risk, and a trail you can follow backwards.
Look at that spread and the switch-flipping instinct makes perfect sense. Tens of thousands of fetches for a single visit is not a fair forage, it is a warren being stripped. The instinct is sound and the implementation is where sites get hurt, because blocking AI crawlers at the zone level is a hammer swung in a dark den, aimed at a shape you cannot quite see.
The crawler table: who is who, and what a block really costs#
Every serious conversation about blocking AI crawlers should start with a table rather than a toggle. Print this one, take it to whoever owns the CDN, and make them say out loud which rows they intend to close. Half the risk evaporates the moment the decision is named crawler by crawler instead of category by category.
| Crawler token | What it is for | What blocking it costs you |
|---|---|---|
| Googlebot | Google Search indexing. Listed as a common crawler that always respects robots.txt for automatic crawls. | Your presence in Google Search. Never block this one. |
| Google-Extended | Controls whether crawled content may train Gemini models and ground responses in Gemini Apps and Vertex AI. | Grounding and training use. Google states it is not a ranking signal and does not impact Search inclusion. |
| Googlebot-News | News crawling. It has no separate user agent string and reuses Googlebot strings. | Google News surfaces. Rarely worth blocking. |
| GoogleOther | A separate common-crawler token with its own robots.txt entry, distinct from Googlebot. | Varies by product. Read the entry before you close it. |
| bingbot | Microsoft's search crawler, verified by reverse DNS to a search.msn.com hostname. | Your presence in Bing, and in surfaces that lean on Bing's index. |
| GPTBot | OpenAI's crawler for making generative foundation models more useful and safe. | Training use of your content. No ChatGPT search penalty on its own. |
| OAI-SearchBot | Surfaces websites in ChatGPT's search features. | Visibility inside ChatGPT search. Block only if you truly want out. |
| ChatGPT-User | Fetches a page when a person asks ChatGPT about it, or via GPT Actions. | Live user-initiated fetches. Blocking makes you unquotable on demand. |
| ClaudeBot | Anthropic's crawler collecting content that may contribute to training. | Training use only. Anthropic documents it separately from search. |
| Claude-SearchBot | Navigates the web to improve search result quality. | Indexing for Claude's search answers. A quiet GEO cost. |
| Applebot-Extended | Named alongside Google-Extended in Cloudflare's managed robots.txt directives. | Apple's AI training use, per Cloudflare's own published wording. |
Never trust the string on its own, because anyone can forge a user agent header. Bing's webmaster team tells site owners to run a reverse DNS lookup on the requesting address, confirm the hostname ends in search.msn.com, then run a forward lookup back to the same address. Bing deliberately publishes no crawl IP ranges, because they change and hardcoded lists end up blocking bingbot by accident.
Google names the same three signals in its crawler overview: the user agent header, the source address, and the reverse DNS hostname. That page also splits Google's fleet into common crawlers, special-case crawlers and user-triggered fetchers, and notes that AdsBot can ignore the global robots.txt rule with publisher permission. Sensible ai bot management seo work starts by knowing which of those three buckets a request sits in.
The failure is silent, and silence is the expensive part#
A 403 is not a polite refusal, it is a deletion order with better manners, and it is how blocking AI crawlers turns into losing pages. Google's documentation on HTTP and network errors states that all 4xx errors except 429 are treated the same way: the crawler tells the next processing system that the content does not exist, and for Search the indexing pipeline removes the URL from the index if it was previously indexed. Read that sentence to whoever owns the edge configuration.
Search Console will tell you eventually, in a report most teams open far too rarely. The page indexing report carries a dedicated status for pages blocked by a forbidden response, and Google's guidance is direct: Googlebot never provides credentials, so a 403 aimed at it means the server is returning that error incorrectly. The remedy Google gives is to admit visitors who are not signed in, or to allow verified Googlebot requests through without a credential challenge. Blocking AI crawlers should never produce that status for a search bot.
Two comfortable myths worth a clean kill#
First, robots.txt is not an index control. Google's introduction to robots.txt says outright that it is not a mechanism for keeping a page out of Google, and a disallowed URL that other sites link to can still appear in results, simply without a description. Second, robots.txt is advisory. The standard itself, RFC 9309, published in September 2022, binds only the crawlers that choose to be bound.
There is a genuinely nasty detail buried in that RFC which almost nobody has read. If robots.txt returns a 5xx, a compliant crawler must assume complete disallow. If it returns a 4xx, the crawler may access any resource on the server. So a blanket edge rule that answers 403 to any unrecognised agent does not tighten your rules, it deletes them for every bot still willing to knock. Blocking AI crawlers badly can leave you looser than blocking nothing at all.
Crawl capacity makes the lag worse, though not quite in the way most people picture it. Google's guidance on managing crawl budget for large sites explains that server errors, rate-limiting signals and slower responses pull the crawl capacity limit down, and that crawl demand differs per crawler while the capacity itself is shared.
A forbidden response is a different beast in that respect, because Google's error documentation states that 4xx codes other than 429 have no effect on crawl rate. So the block does not slow the crawl politely, it strips the pages out while the crawler keeps knocking. Documentation on crawl budget is worth a careful read before anyone starts throttling bots at the edge, because a site that crawls slowly recovers slowly.
A block is a business decision wearing a network setting's clothes. Someone in marketing should be in the room when it is made.
A five-step audit to run before 15 September 2026#
None of this needs a replatform, a retainer or a heroic sprint. It needs one afternoon, one shared document, and one person willing to ask an awkward question in a channel where nobody usually asks. Run the audit now, while blocking AI crawlers is a choice you are making rather than a default you inherited on a date somebody else set.
Request your sitemap, homepage and three money pages as Googlebot and as bingbot from outside your network. Anything other than a 200 is the whole story, and a googlebot 403 error on a sitemap is a five-alarm finding, not a curiosity.
List every live bot rule on the zone, its owner, its date and its intention. Rules with no named owner get closed or claimed. Undocumented edge rules are the undergrowth this problem hides in.
Rewrite blanket blocks as named robots.txt disallows: Google-Extended, Applebot-Extended, GPTBot, ClaudeBot. Explicitly allow Googlebot, bingbot and OAI-SearchBot so nobody has to guess your intention later.
Open your CDN security settings and confirm what happens on 15 September 2026. Decide the opt-out deliberately, record who decided, and diarise a review for the week after the date lands.
Alert on non-200 responses to verified search-engine user agents, and put the page indexing report on a monthly calendar invitation with a named owner. Blocking AI crawlers without monitoring is a snare set in your own path.
The reporting layer matters as much as the rules. If impressions slide and nobody can say whether the cause was an algorithm, a competitor or a checkbox, the argument in the review meeting will be settled by whoever speaks with most conviction rather than most evidence. That is how good sites lose quarters. Our work on AI visibility and citation reporting exists for exactly that reason.
There is a strategic reading here too, and it is not the doom-laden one. Being visible to AI systems is now part of the same discipline as being visible in search, which is why we treat them as one brief in SEO and GEO services rather than two competing workstreams. Blocking AI crawlers wholesale opts you out of the answer layer as surely as a noindex tag opts you out of the results page, and the pieces on SEO versus GEO and the hidden costs of opting out make the same point from the other direction.
Regulated categories carry the extra weight, as ever. An operator in iGaming marketing or FinTech marketing that vanishes from the index does not merely lose traffic, it loses the ability to correct what a model says about its licence while somebody else's page fills the gap. Sound content marketing assumes the content can be reached. Check the gate before you trust the trail.
The fox lesson, then, and it costs nothing to learn twice. Do not close the whole hedgerow. Find the gap the quarry actually uses, close that one, mark it on a map somebody else can read, and come back in September to check the gate still swings the way you left it.
"On September 15, Cloudflare will start cutting off Google traffic for millions of its smallest customers by default. And those customers won't leave. They've been begging for this...
Frequently asked questions#
Does blocking AI crawlers stop Google from indexing my site?
It can. Google-Extended is a separate token, and disallowing it does not affect Search. A blanket edge rule that returns 403 to unrecognised agents is different: Google treats a 403 as content that does not exist and removes previously indexed URLs, so the damage appears as pages quietly leaving the index rather than as an error anyone gets alerted to.
What actually changes on 15 September 2026?
Cloudflare has said its defaults will block Training and Agent crawlers on pages carrying ads for new domains, while Search stays allowed. Because multi-purpose crawlers such as Googlebot, Applebot and BingBot are judged by every purpose they serve, customers who block Training block those bots too. An opt-out sits in security settings before the date.
How do I block AI training without losing search traffic?
Name the tokens instead of the category. Disallow Google-Extended, Applebot-Extended, GPTBot and ClaudeBot in robots.txt, and explicitly allow Googlebot, bingbot and OAI-SearchBot. That is blocking AI crawlers with a scalpel rather than a hammer, and it keeps every indexing crawler you actually depend on.
Is blocking AI crawlers ever the right call?
Often, yes. If your content is your product, blocking AI crawlers that only train models is a defensible commercial decision, and vendors now publish separate tokens precisely so you can make it. The failure mode is not the intention, it is the blunt instrument: category-level blocks that sweep up multi-purpose search crawlers you never meant to touch.
How do I check a crawler really is Googlebot or bingbot?
Verify the source rather than the string, because user agents are trivially forged. Run a reverse DNS lookup on the requesting address, confirm the hostname belongs to the search engine, then run a forward lookup back to the same address. Bing points at hostnames ending in search.msn.com; Google lists the user agent, the address and the reverse DNS hostname as its three signals.
Will robots.txt keep my pages out of AI answers?
Only for crawlers that honour it. The robots exclusion standard is voluntary, and Google states that robots.txt is not a mechanism for keeping a page out of Search at all, since a disallowed URL linked from elsewhere can still appear without a description. For genuine exclusion you need noindex, authentication or a rule enforced at the edge.
Who should own the decision about blocking AI crawlers?
Marketing, SEO and engineering jointly, with one named owner and a written record. The setting lives in infrastructure but the consequence lands on revenue, so a change should carry a date, an intention and a rollback note. Review it on a schedule rather than the first time somebody notices impressions sliding.
Read more on this topic#
AI Brand Visibility: what 3,960 model answers revealed
Familiarity beat relevance in nearly four thousand model answers. What that means for brands nobody has heard of yet.
Read the studyAI Advertising Agents and the Ads MCP server
Agents are starting to buy media. The protocol underneath them is worth understanding before your budget meets one.
Read the pieceThe opt-out button that quietly removes you from the news
The other switch with an invisible price tag, and why the Search Console opt-out costs more than it advertises.
Read the pieceSEO vs GEO: why your best pages miss AI citations entirely
Ranking and being quoted are two different jobs. Here is what separates a page that gets cited from one that gets summarised.
Read the piece
Want the gate checked before September?
folkfox audits crawler access, edge rules and AI visibility together, so blocking AI crawlers stays a decision you made rather than a default you inherited, with reporting that flags the moment a setting starts costing you the index.