Skip to main content

folkfox

Skip to main content
Skip to content
SEO and GEO

AI Crawl Control now lets you refuse the trainer and keep the search

On 15 September 2026, Cloudflare stopped forcing regulated sites to choose between Google and their own content: one setting now says no to training and yes to search.

Quick answerAI Crawl Control's new Disallow AI Training setting writes no-training rules for Google-Extended and Applebot-Extended while Googlebot, Applebot and Bingbot keep crawling for search. Microsoft's robots.txt support is targeted for early 2027.
Section 01

What AI Crawl Control changed on 15 September#

A fox at a hedgerow gate has two clumsy options: bar it against every traveller, or leave it swinging in the wind. For years that was the whole menu for a site owner who disliked AI training but lived on Google traffic, because the biggest crawlers do two jobs at once. Cloudflare's own account of the trap, in its 15 September announcement on accountable mixed-use crawlers, is that refusing one purpose meant refusing the other.

How many sites refuse training
Waffle chart: 17 percent of Cloudflare sites use some mechanism to block AI training, the ai crawl control default this article examines17% of Cloudflare sites use some mechanismto block AI training
Waffle chart: 17 percent of Cloudflare sites use some mechanism to block AI training, the ai crawl control default this article examines
ItemValue
17% of Cloudflare sites use some mechanism17% of Cloudflare sites use some mechanism
to block AI trainingto block AI training
Seventeen in every hundred Cloudflare sites block AI training in some way, while fewer than one in a hundred block search bots, per Cloudflare's announcement.

The new answer is a setting called Disallow AI Training. Cloudflare frames the release as changes to Bot Management and AI Crawl Control, and the switch itself sits in a zone's Security Settings. What it does is small and sharp: it writes a no-training rule into robots.txt for Google-Extended and Applebot-Extended, so Googlebot, Applebot and Bingbot keep crawling for search. Search Engine Journal's report notes that this reverses the plan Cloudflare floated in July, when blocking training would have shut those three search crawlers out as well.

That reversal matters for anyone who runs a regulated site and hunts quarry across the search thicket, because the old menu made the safe-sounding choice the costly one. Under the new AI Crawl Control arrangement, refusing training and staying findable stop being opposites, at least for the three crawlers that matter most to a search-led business.

The trade Cloudflare struck: access for accountability#

The reversal was not a gift. Cloudflare says a crawler earns its "Accountable" label by meeting four tests: a way for site owners to opt out of AI training through robots.txt or a similar standard; a way to opt out of AI summaries, set with the operator directly; URL-level visibility into which pages were made available for training, alongside search metrics; and an assurance that opting out of training will not touch traditional search results.

Cloudflare labels Google, Apple and Microsoft Accountable, and its post says Microsoft still needs to add robots.txt support for a no-training preference, which it targets for early 2027.

Then comes the number that keeps the whole story honest. Cloudflare reports that less than 1 percent of its sites choose to block search bots, and that 17 percent enable some mechanism to block training. Refusing the trainer is a real minority habit, and refusing the searcher is close to a rounding error.

Section 02

What Cloudflare AI Crawl Control does, and what it leaves alone#

Read the small print of the new arrangement before you tick it, and before you brush past the dashboard's own name for it. Cloudflare's developer documentation describes AI Crawl Control as a way to monitor and control how AI services reach your content, with visibility of which crawlers arrive, allow-or-block policies for each, and monitoring of who obeys robots.txt. Cloudflare AI Crawl Control is therefore a dashboard and a policy layer in the den of your security settings, and the training setting is best read as a robots.txt writer with a label attached, not a wall.

Underneath, Bot Preference Sync keeps robots.txt in step with the search, agent and training policy you choose. Cloudflare says in its announcement that it is deprecating the older cloudflare block ai bots switch in favour of those three finer controls, and retiring Managed Robots.txt in favour of the sync. Search Engine Journal adds that sites using the older options migrate automatically to Disallow AI Training.

Google-Extended is the lever most people mean when they say opt out of Gemini. Google describes it as a standalone product token that publishers use to manage whether content Google crawls may train future Gemini models and ground answers in Gemini apps. It has no user agent string of its own, and Google's crawler documentation says it does not affect inclusion in Google Search and is not a ranking signal. That is why the scent of this setting is safe: you are asking Google to stop one use of a crawl it keeps doing anyway.

Watercolour fox at an open garden gate, one paw on a closed book, a lantern lit above: AI Crawl Control refusing training while search stays open
The gate stays open for search; only the book is closed.

Apple's Applebot-Extended is gentler still. Apple's support page says it does not crawl webpages at all; it only records your wish that content not be used to train Apple's general-purpose foundation models. Pages that disallow it can still be included in search results, and Apple says the primary Applebot keeps powering search across Spotlight, Siri and Safari. So applebot-extended costs you nothing in findability by Apple's own account, and google-extended costs you nothing by Google's.

Four things Disallow AI Training leaves untouched#

First, it is a request, not a wall. Cloudflare's post is candid that a robots.txt directive alone cannot identify who is crawling, work out why, or stop a crawler that ignores it, and the Bot Preference Sync post adds that crawlers which do not offer transparency are still blocked when you disallow training. For the Accountable three, the bargain rests on their word.

Second, it does not decide your place in AI Overviews or AI Mode. Google's Search generative AI control in Search Console is a separate switch that includes or excludes your content from AI Overviews, AI Mode and generative features in Discover, and Google states that it does not affect AI training, which is Google-Extended's job. Snippet controls such as nosnippet, listed in Google's AI features guidance, are a third lever. Three levers, three different jobs.

Third, it does not cover Bing yet. Microsoft's 2023 announcement describes NOARCHIVE as keeping content out of Bing Chat answers and out of training its generative foundation models, while NOCACHE lets content appear in answers as a URL, title and snippet that may still be used for training; either way the page stays in Bing search. Until Microsoft ships robots.txt support, a Bing-side refusal is a meta tag, and the trade is training against citation.

Fourth, none of the pages I read says a no-training preference reaches back into models already trained. Treat it as a rule for what happens next, not an eraser for what happened before.

Google and Apple offer a robots.txt opt-out that leaves search intact today, while Microsoft's equivalent is still a promise for early 2027.
OperatorNo-training leverEffect on search, per the vendorStatus
GoogleGoogle-Extended token in robots.txtNot a ranking signal, no effect on inclusion in SearchLive
AppleApplebot-Extended in robots.txtDisallowed pages can still appear in searchLive
MicrosoftRobots.txt support planned; NOARCHIVE or NOCACHE meta tags meanwhileTagged pages still appear in Bing searchRobots.txt targeted early 2027
  • GoogleGoogle-Extended token in robots.txtNot a ranking signal, no effect on inclusion in Search Live
  • AppleApplebot-Extended in robots.txtDisallowed pages can still appear in search Live
  • MicrosoftRobots.txt support planned; NOARCHIVE or NOCACHE meta tags meanwhileTagged pages still appear in Bing search Robots.txt targeted early 2027
Section 03

Why regulated sites now face a decision, not a default#

Healthcare, fintech, iGaming and cybersecurity share a habit: their pages carry statements someone can be held to. A licence number, a rate, a clinical scope, an advisory. For those sites, allow everything and block everything were both reckless defaults, and the honest position sat in the undergrowth between them. AI Crawl Control now lets that middle path be a single setting, which turns a default into a decision, and decisions need owners.

Three separate facts about AI and search
US adults read AI summaries
60%
Sites blocking AI training
17%
Sites blocking search bots
under 1%
Six in ten US adults say they read AI summaries, yet 17 in 100 Cloudflare sites refuse training and under 1 in 100 refuse search: three populations, so read the order, not the gaps. "Under 1 percent" is drawn as a 1 percent ceiling. Sources: Pew Research Center and Cloudflare.

Read those bars as three separate facts, not one league table. The Pew survey, fielded 17 to 23 February 2026 and published on 17 June, found 60 percent of US adults say they read AI summaries at the top of search results. That is where your words now get met by a stranger, and Disallow AI Training does not govern it. The bar for Cloudflare's 17 percent measures a different population entirely, so the comparison is about order of magnitude, not arithmetic.

Which leaves three questions that no AI Crawl Control default can answer for you. Do you refuse training? Do you take part in the summary? And do you want to be paid, or at least credited, for the use? The first is now cheap. The second is a Search Console choice with a few days' lag, per Google's help page. The third is commercial, and it lives outside robots.txt, as we argued in Three files tell you the whole strategy and as the licensing side of the same coin, the case for direct AI feeds, explores.

Refusing the trainer and keeping the search is now one switch; deciding about the summary is still a separate, deliberate act.
folkfox, on the choice AI Crawl Control leaves with you

Four regulated lanes, four different first questions#

In healthcare, the first question is faithfulness. A page that states who a service is for, or what it does not do, is a statement a regulator might read, so the priority is whether a summary paraphrases it accurately rather than whether a model trained on it. Training refusal is the cheap half of the job for healthcare marketing teams; the summary is the costly half.

In fintech, the first question is disclosure. Rates, fees and risk warnings only work if they travel with the claim. A robots.txt rule cannot make a summary carry the warning, and neither can a training opt-out, so the practical guardrail is writing the risk sentence so it stands alone, which is the discipline behind fintech marketing that survives any crawler policy.

In iGaming, the first question is the licence sentence, and we have already argued in MGA Compliance is AI Social Proof that a stated, verifiable licence is a citation asset. That makes refusing training an easy call for iGaming marketing, and staying inside the summary a deliberate one.

In cybersecurity, the first question runs the other way. Advisories and research exist to be shared, so an advisory may want every reader, human or model, while gated research wants a closed gate. For cybersecurity marketing, the answer is per section of the site, not per domain, and a single robots.txt line rarely encodes that nuance.

Section 04

What to check first: the zone, robots.txt and your licence#

Now the practical prowl through the AI Crawl Control dashboard. The order matters, because the costliest mistake is the quiet one: a search crawler blocked without anyone noticing. Cloudflare and Search Engine Journal both report that the Block setting now applies to mixed-use crawlers, so selecting Block halts Googlebot, Applebot and Bingbot for search as well as training. Disallow AI Training is the setting that keeps them in, and the whole check begins with confirming which one your zone is actually wearing.

Five checks, in order
Read the zone

Open Security Settings and read the current Search, Training and Agent values in AI Crawl Control. Confirm Search is Allow.

Read robots.txt

Look for Google-Extended and Applebot-Extended lines, and for any leftover lines from the older managed file.

Watch the crawl

Check Search Console for crawl and sitemap-fetch errors after the migration, and Bing Webmaster Tools too.

Decide the summary

Choose separately whether to stay in AI Overviews and AI Mode using Google's Search generative AI control.

Write it down

Record your training, summary and licensing positions with an owner and a date, and diarise Microsoft for early 2027.

The first check deserves a second look, because Search Engine Journal reports that existing Block AI Bots users migrate automatically. Do not trust your memory of what you set. Trust what the dashboard now says, and compare it with what robots.txt now publishes, since Bot Preference Sync writes that file on your behalf and a stale copy from a plugin or a CDN rule can contradict it.

The second check is the licensing position, and no AI Crawl Control toggle can set it. A robots.txt line is a preference, and a contract is a contract. If you sell, or plan to sell, access to your content, a blanket refusal today shapes tomorrow's conversation, and refusing training for free while licensing it to a partner is a story your own counsel should approve first. Nothing in Cloudflare's setting, Google's token or Apple's token creates or removes a licence.

The third check is for the search side of the house. If your crawl stats drop after a settings change, the cause is nearly always at the edge, not in Google, and the fix belongs to whoever owns the Cloudflare dashboard. That is exactly the boundary where folkfox's SEO and GEO services work with an in-house engineer: we read the crawl, they own the switch.

Section 05

No llms.txt and no GEO hack required#

There is a louder story running alongside this one, and it is worth answering plainly. Vendors keep selling tricks to appear in Google's generative features. Google's guide to generative AI features, updated 10 July 2026, says you need no new machine-readable files, AI text files, markup or Markdown to appear in Google Search, including its generative capabilities, because Google Search does not use them. It adds that llms.txt is fine to keep for other systems, and that it neither helps nor harms your visibility in Google.

The usage data agrees. Ahrefs looked at 137,210 domains in its Web Analytics that had traffic in May 2026: 28 percent published an llms.txt file, and 97 percent of those files received zero requests in the month. Ahrefs also found that AI tools never go looking for llms.txt files that are not there. So the house line at folkfox holds: no llms.txt and no AI-specific markup is needed to appear in Google's generative features, and we will not sell you a hack that Google says it ignores. Our AI visibility work and AI consultancy start from the same premise.

What does move the needle is dull and durable: pages a crawler can fetch, sentences a summary can quote without distorting, and a policy someone owns. That is the thread through AI search visibility needs an earned citation trail and AI overview tracking still beats waiting, and it is why a training opt-out belongs in the same meeting as your citation reporting.

The standard to watch, and the date to diarise#

Two clocks are ticking. The first is Microsoft's, with robots.txt support for a no-training preference targeted for early 2027 by Cloudflare's account.

The second clock is the IETF's AI Preferences working group, whose charter commits it to a standard vocabulary for expressing AI preferences and to ways of attaching them, including through the Robots Exclusion Protocol and HTTP headers. It lists 31 August 2026 as the target for its vocabulary and attachment drafts, and I did not confirm whether those milestones landed. Until a standard settles, a robots.txt token per crawler is the working language, and AI Crawl Control is the tool that writes it for you. Follow the trail one crawler at a time and you will outfox both the hype and the hazard.

So the decision is yours, and it is smaller than the headlines suggest. Refuse the trainer with one setting, keep the search gate open, decide the summary on purpose, and write the licensing position down. On 15 September 2026 AI Crawl Control made the first of those cheap. What will Microsoft's early 2027 promise make cheap next?

Questions

Frequently asked questions#

What is AI Crawl Control?

AI Crawl Control is Cloudflare's product for monitoring and controlling how AI services reach your content, with per-crawler visibility, allow or block policies and robots.txt compliance monitoring. On 15 September 2026 Cloudflare added a Disallow AI Training setting to the same family of controls.

What does Cloudflare's Disallow AI Training setting do?

It publishes a no-training preference in robots.txt, for Google-Extended and Applebot-Extended, while Googlebot, Applebot and Bingbot continue to crawl for search. It is a preference that Accountable crawlers have agreed to honour, not a technical wall, and it sits within AI Crawl Control.

Does blocking Google-Extended hurt my Google rankings?

Google says it does not. Its crawler documentation states that Google-Extended does not affect inclusion in Google Search and is not a ranking signal. It is a separate token, with no user agent of its own, that governs Gemini training and grounding.

Should I disallow Applebot-Extended?

That is a licensing and policy decision, not an SEO one. Apple says Applebot-Extended does not crawl, and pages that disallow it can still appear in search across Spotlight, Siri and Safari. Disallowing it only opts your content out of training Apple's general-purpose models.

What happens to the Cloudflare block AI bots setting?

Cloudflare is deprecating the old Block AI Bots switch in favour of separate Search, Training and Agent controls, and retiring Managed Robots.txt in favour of Bot Preference Sync. Search Engine Journal reports that existing sites migrate automatically to Disallow AI Training.

Do I need an llms.txt file to appear in Google's AI features?

No. Google says you do not need new machine-readable files, AI text files, markup or Markdown to appear in Google Search, including its generative features, because Google Search does not use them. Ahrefs found 97 percent of llms.txt files received no requests in May 2026.

Does Microsoft support a robots.txt opt-out for AI training yet?

Not yet. Cloudflare says Microsoft still needs to add robots.txt support for a no-training preference, and targets early 2027. Until then, Bing offers NOARCHIVE and NOCACHE meta tags, with a trade-off between training and appearing in Bing Chat answers.

Keep reading

Read more on this topic#

Ready to refuse the trainer and keep the search?

Your Cloudflare policy, robots.txt and crawl logs get read together at folkfox, then written into the training, summary and licensing positions your regulated site can defend.

Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.