Skip to main content

folkfox

Skip to main content
Skip to content
SEO AND GEO

Your robots.txt is a request , and the fetchers know it

A file in your web root has been carrying more weight than it was ever designed to hold. For most sites that is a nuisance. For anyone with jurisdiction-gated pages, it is a compliance question wearing a technical costume.

Quick answerAI crawlers are automated fetchers that retrieve pages for AI training, search and user-initiated answers. Some documented fetchers do not apply robots.txt at all, so a disallow line is a request rather than a lock.
SECTION 01

What robots.txt was always for#

ai crawlers

Every fox knows the difference between a fence and a sign. One stops you. The other tells you the farmer would rather you did not. The robots exclusion protocol has always been the second thing, and the standard that formalised it says so in plain language. That distinction is the whole reason ai crawlers can behave the way they currently do.

The protocol was written up properly in 2022 as RFC 9309, on the IETF standards track. It is worth reading once, because it settles arguments. The document is explicit that these rules are not a form of access control, and it sets out matching behaviour that surprises people: "If no match is found amongst the rules in a group for a matching user-agent or there are no rules in the group, the URI is allowed".

In other words, silence is permission. A file that does not name a fetcher has not blocked it, which is why lists of ai crawlers go stale faster than anyone maintains them. And even a well-formed disallow only binds a crawler that has chosen to read the file, matched itself to a group, and decided to obey. Three separate acts of goodwill, none of them enforceable.

Cached goodwill, at that#

The same standard tells crawlers they "SHOULD NOT use the cached version for more than 24 hours", which means a rule you added this morning may not bind ai crawlers until tomorrow. Anyone who blocks a path in response to an incident should assume a day of overlap rather than an instant seal.

None of this is a scandal or a loophole. It is the design, working as designed, in an era that is asking far more of it than its authors intended. Two of RFC 9309's four authors work at Google, and the document reads like people describing a convention they wished to preserve rather than a wall they were building.

SECTION 02

What the ai crawlers themselves say#

Here is where the trail gets interesting, because the labs running ai crawlers have documented their own behaviour and the documents do not agree with each other. This is not leaked information or inference. It is published policy, and it is the strongest evidence available on this question.

OpenAI documents several bots on its bots documentation, and separates them by purpose. GPTBot gathers training data, and the page states that "Disallowing GPTBot indicates a site's content should not be used in training generative AI foundation models". OAI-SearchBot handles search indexing. Then there is ChatGPT-User, the fetcher that runs when a person asks ChatGPT to go and look at something, and OpenAI's own wording on it is the line that matters: "Because these actions are initiated by a user, robots.txt rules may not apply."

Perplexity draws the same distinction and is, if anything, blunter. Its crawler documentation describes PerplexityBot as the compliant search crawler, then says of the user-triggered fetcher: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules."

Anthropic takes the opposite position on precisely the same category of traffic. Its crawler documentation lists ClaudeBot for training, Claude-User for user-initiated fetches and Claude-SearchBot for search indexing, and states that its bots respect "do not crawl" signals through standard robots.txt directives. That statement covers the user-initiated fetcher too, which is the meaningful difference.

Taken from each company's own published crawler documentation. The disagreement is about user-initiated fetches, not about training crawlers.
FetcherPurposeDocumented robots.txt posture
GPTBotModel trainingRespects robots.txt
OAI-SearchBotChatGPT search indexingRespects robots.txt
ChatGPT-UserUser-initiated fetchRules "may not apply"
PerplexityBotSearch indexingRespects robots.txt
Perplexity-UserUser-initiated fetch"Generally ignores" the rules
ClaudeBotModel trainingRespects robots.txt
Claude-UserUser-initiated fetchRespects robots.txt

Read that table as a fox reads a hedgerow: the gaps are the story. Every operator of ai crawlers agrees that a training crawler should obey. They disagree about whether a machine acting on a human's instruction inherits the human's right to visit a page. That is a genuinely hard question, and it has been answered three different ways by three companies.

SECTION 03

How much actually leaks through#

Policy is one thing, measurement is another, and this is where the evidence gets thinner and needs handling with care rather than enthusiasm.

The strongest independent work is academic. A large-scale empirical study published as arXiv preprint 2505.21733 tracked "130 self-declared bots (and many anonymous ones) over 40 days" using institutional web logs alongside controlled robots.txt experiments. Its finding is stated without hedging: "bots are less likely to comply with stricter robots.txt directives, and that certain categories of bots, including AI search crawlers, rarely check robots.txt at all". The authors' conclusion is the sentence to take to a stakeholder meeting: "relying on robots.txt files to prevent unwanted scraping is risky".

Vendor measurement of ai crawlers points the same way, though it is harder to pin down than it should be. The licensing vendor TollBit has published bot reports for several quarters, and its earlier public editions carried a bypass rate in the low teens as a percentage of AI bot visits. Those earlier pages now redirect, so this article does not quote a figure from them.

The 2026 figure, and why it carries an asterisk#

A newer TollBit number has been circulating this month: roughly 15% of identified AI fetching agents in Europe reaching URLs the site had disallowed, with ChatGPT-User named as both the most blocked and the most frequent offender. That was reported by PPC Land.

Being straight about provenance, because it matters more than the number does: the 2026 report is titled on TollBit's State of the Bots index, whose fourth section is named for robots.txt and the bypassing problem, but the report itself sits behind a download form. The 15% could not be checked against TollBit's own publication for this article, so it is reported here as a trade-press figure and nothing more. The academic study above and the crawler documentation below are what this piece actually rests on.

Documented posture of seven named ai crawlers
Donut chart showing five of seven documented ai crawlers respect robots.txt while two state the rules may not applyDocumented as respecting robots.txt: 71%Documented as not bound by it: 29%2 of 7
Documented as respecting robots.txt 71%Documented as not bound by it 29%
Counted from the seven fetchers in the table above, each classified by its own operator's published documentation. Two are documented as not bound by robots.txt, and both are user-initiated fetchers.
The academic sample, in numbers

Self-declared bots tracked

130

Plus an undisclosed number of anonymous ones, over institutional logs.

Days observed

40

Log analysis paired with controlled robots.txt experiments.

Fetchers documented as unbound

2

ChatGPT-User and Perplexity-User, per their operators own published pages.

SECTION 04

Why this is a compliance problem, not an SEO one#

For a publisher, a leaking robots.txt is a revenue argument. For a licensed operator it is something sharper, and this is the part that tends to get missed in the general commentary.

Regulated businesses routinely publish pages that are lawful in one jurisdiction and unlawful in another. Bonus terms a Maltese licence permits and a British one does not. Product wording cleared for a professional investor and forbidden to a retail one. Treatment claims a prescriber may read and a patient may not. The usual control is geographic gating plus a disallow line aimed at ai crawlers, on the assumption that the second one holds.

It does not reliably hold. A user-initiated fetcher that does not apply your robots.txt can retrieve a page written for one market and summarise it for a reader in another, and no server log of yours will show a human visiting from the wrong place. The exposure is not that your content was copied. It is that a restricted claim was repeated to an audience your licence does not cover.

The disallow line is the control

Gated pages are listed in robots.txt, the compliance file records that they are blocked from crawlers, and nobody revisits the assumption because nothing has visibly gone wrong.

The disallow line is a preference

Access control lives at the server: authentication, IP rules, or simply not publishing the page. robots.txt records an intention, and intentions do not appear in enforcement correspondence.

The fox's instinct applies neatly here. If you would be uncomfortable seeing a sentence quoted back to you by a stranger in another country, that sentence does not belong on a page whose only protection from ai crawlers is a text file asking politely.

SECTION 05

Measure first, then decide whether to block ai crawlers#

The reflex when people meet this evidence is to reach for a wholesale block. Resist it for a fortnight, because the decision splits into two questions that get answered together and should not be.

The first question is exposure: which pages could be fetched that should not be. The second is ai visibility: whether being retrievable by answer engines is worth anything to you commercially. Those have different answers for a licensed casino's bonus terms and for its careers page, and a single blanket rule in a single file cannot express that difference well.

Tooling has arrived, and it measures the easy half#

Microsoft added scrape-to-referral reporting to Clarity, ranking AI operators by crawler requests against human sessions, as covered by Search Engine Land. It turns a vibe into a ratio, which is genuinely useful for the argument about whether a given operator gives anything back.

Any ai visibility checker on the market measures roughly the same easy half: what got fetched, what got cited, what sent a visitor. None of them can tell you whether a page that leaked contained a claim your regulator would object to. That judgement stays human, and it belongs with whoever owns the compliance file rather than whoever owns the analytics login.

Scale is not hypothetical either. Mediavine created a publisher advocacy role after Google Search and Discover referrals across its network fell by around a third, reported by PPC Land. In France, the main press publishers association filed a competition complaint alleging AI summaries cut traffic substantially, also covered by PPC Land.

Where blocking actually helps
Jurisdiction-gated compliance pages
block
Paywalled or licensable archives
block
Proprietary research and data
consider
Service and product pages
leave open
Careers, press and brand pages
leave open
Directional ranking from folkfox client work, not a benchmark. Blocking pays where content is restricted or licensable, and costs where discovery is the point.
The order that keeps you honest
Inventory the restricted

List every page whose lawfulness depends on who is reading it. Jurisdiction, licence class, audience type. This list is usually shorter than people fear and never empty.

Read your own logs

Pull the user agents actually hitting those paths for a full month. Name them. You cannot block what you have not named, because an unnamed fetcher is an allowed one.

Move the real controls

Anything genuinely restricted moves behind authentication, IP rules or off the public site entirely. robots.txt stays as a statement of intent, not as the lock.

Choose visibility deliberately

For everything else, decide whether retrieval is worth having. Most service pages benefit from being fetchable. Say so explicitly rather than blocking by default.

Re-read the documentation quarterly

These policies change without announcement. The posture table in this piece was accurate on the day it was written and is not a permanent fact.

A disallow line records what you wanted. Access control records what happened. Only one of those turns up in an enforcement letter.
folkfox, on robots.txt and regulated content

If you want the inventory built and the log review run rather than described, that is folkfox SEO and GEO services, applied with the same care across iGaming marketing, FinTech marketing and healthcare marketing, the three categories where a leaked sentence costs most.

Related reading from the folkfox den: Cloudflare's free AI visibility scoreboard covers the measurement side, and AI Overviews SEO covers what happens once you have decided to be found.

Questions

Frequently asked questions#

Should I block AI crawlers on my whole site?

Rarely. Blocking pays on jurisdiction-gated, paywalled or licensable pages, and costs you on service and brand pages where being retrievable is the point. Decide page by page rather than issuing one blanket rule, and move genuinely restricted content behind real access control instead.

What are AI crawlers actually doing on my site?

Three different jobs, usually. Training crawlers gather text for model building, search crawlers index pages so they can be cited in answers, and user-initiated fetchers retrieve a specific page because a person asked an assistant to look at it. Each has a different documented policy.

Does robots.txt legally stop a crawler?

No. RFC 9309 describes a voluntary convention, not an access control mechanism. A crawler must choose to read the file, match itself to a group and obey it. Where no rule matches a fetcher, the standard says the page is allowed.

Which AI fetchers ignore robots.txt?

By their own documentation, OpenAI's ChatGPT-User states robots.txt rules may not apply because a person initiated the request, and Perplexity-User generally ignores the rules for the same reason. Anthropic documents that its user-initiated fetcher does respect robots.txt directives.

Cloudflare's 22 Aug 2026 Bot Preference Sync launch post: robots.txt is a request, not a lock, and AI crawlers treat it accordingly.

Can an AI visibility checker tell me if I have a compliance problem?

No. Those tools measure what was fetched, what was cited and what sent traffic. They cannot judge whether a leaked page carried a claim that is unlawful in the reader's jurisdiction. That assessment stays with whoever owns your compliance file.

How quickly does a new robots.txt rule take effect?

Not instantly. RFC 9309 tells crawlers they should not use a cached copy for more than 24 hours, so assume up to a day of overlap after you add a rule. If a page needs to stop being reachable immediately, remove it or put it behind authentication.

Keep reading

Read more on this topic#

Not sure what your logs are already showing?

folkfox audits crawler exposure for businesses whose pages are lawful in one market and awkward in another, then builds the visibility decision on evidence instead of instinct.