Cloudflare adds Disallow AI Training to keep crawlers in search

Cloudflare announced a new setting called Disallow AI Training, aimed at a specific problem: some of the largest crawlers on the internet, run by Apple, Google, and Microsoft, are mixed-use crawlers that both build a search index and train AI models on the same visit. Until now, a site owner who wanted to keep that crawler out of AI training had only one lever available: block the crawler outright, and lose search visibility along with it. The new setting is a second lever. Turn it on for a domain, and Cloudflare's Bot Preference Sync publishes a Disallow: directive for AI training in that site's robots.txt, while the same crawler, if it belongs to an operator Cloudflare designates Accountable, keeps crawling the site for search.
Cloudflare frames the change around its own traffic data: less than 1% of Cloudflare sites choose to block Search bots at all, but 17% of sites enable some mechanism to block AI training, a gap the company reads as evidence that site owners overwhelmingly want to be found while a meaningful share of them do not want their content used to train models. Cloudflare also argues that a robots.txt entry by itself cannot enforce that preference: anyone can publish one, but it cannot identify who is actually crawling, establish why they are crawling, or stop an operator who simply ignores it. Cloudflare's pitch is that its network can do what a text file cannot, by publishing the preference, identifying and classifying each crawler, blocking the ones that ignore it, and then reporting what each operator actually does on its public Radar dashboard.
The Accountable designation grew out of conversations Cloudflare says it has been having directly with crawler operators since July. According to Cloudflare, almost all of the operators it talked to agreed that site owners should have control and transparency over how their content is used, and reassurance that their stated choices will actually be respected. To carry that agreement into a public label, Cloudflare set four requirements a bot operator must meet, or commit to a timeline for meeting, to be called Accountable: a mechanism for site owners to opt out of AI training, through robots.txt or a similar standard; a mechanism to opt out of AI summaries, set directly with the operator today and, from next year, through Cloudflare itself; URL-level visibility into which pages were made available for training, paired with metrics on how content performed in search; and assurance that opting out of training will not hurt a site's traditional search ranking. Applebot, Bingbot, and Googlebot, the mixed-use crawlers run by Apple, Google, and Microsoft, all qualify. Separately, Cloudflare also categorizes the crawlers run by Amazon, Anthropic, Meta, and OpenAI as Accountable, for a simpler reason: those four companies already run separate Search and Training crawlers, so Cloudflare can block the Training crawler without ever touching search.
The three mixed-use crawlers do not all currently offer the same capabilities. Applebot lets a site opt out of training today by adding a Disallow rule for "Applebot-Extended" to robots.txt, and a site can express an AI Summaries preference through the nosnippet directive in its page HTML, or exclude paywalled content from generative output by labeling it as such; Apple does not yet offer URL-level inspection of what was used for training, though Cloudflare says Apple has shared details of an in-progress solution targeted for next year, and Apple has stated that disallowing training does not affect search ranking. Googlebot takes opt-outs through a Disallow rule for "Google-Extended," gives site owners a toggle in its webmaster portal to exclude content from generative search results, and already reports metrics on both search and AI summary results; Google is working on further URL-level transparency tools tied to Google-Extended that it expects to ship in the weeks to come, and it too says disallowing Google-Extended does not affect ranking. Bingbot offers granular controls and transparency through Microsoft's Webmaster Tools, but its robots.txt-level opt-out is not built yet: Microsoft is targeting early 2027 for a mechanism that respects a domain-wide "no training" preference in robots.txt. Until then, a Bing opt-out still runs through the NOARCHIVE meta tag, or Cloudflare's Block URLs and Content Removal tools, and selecting Disallow AI Training in Cloudflare will not by itself signal a no-training preference to Bing. Microsoft has stated that using NOARCHIVE does not affect search ranking.
The practical settings change on September 15. Before that date, Block and Block on pages with ads did not touch mixed-use crawlers, specifically to avoid collateral damage to search; from that date, both options apply to every training crawler, mixed-use included, so choosing Block now removes Applebot, Bingbot, and Googlebot from a site entirely, search as well as training. The older, blunter "Block AI Bots" toggle is deprecated in favor of the separate Search, Training, and Agent controls, and Managed Robots.txt is deprecated in favor of Bot Preference Sync, with customers on the old system migrated automatically. Cloudflare says almost no existing customer needs to do anything: current settings carry over as they are, and a domain that previously set Training to Block or Block on pages with ads is moved to Disallow AI Training, which Cloudflare says preserves what that earlier choice actually did. Disallow AI Training also becomes part of the recommended default configuration for certain new domains. From September 15, a newly onboarded domain is offered one of two preset configurations, chosen by whether the site earns money from advertising; Cloudflare's reasoning is that ad revenue depends on a human seeing the page, an AI training pass replaces that visit with an answer, and an agent fetch leaves nobody there to see an ad either, so the preset for ad-funded sites is the more restrictive one. Either preset can be changed during onboarding or later.
Disallow AI Training is available only as a Training-level setting, not for Search or Agent. Cloudflare says an ads-only exclusion cannot be published the same way, since the list of which pages currently carry ads is too large and changes too often to enumerate in a robots.txt file. Agents, meaning user-directed tools that fetch a page on a person's behalf such as chat-fetch bots and browser-use agents, get no Disallow option yet either: Cloudflare argues agents do not create the same search-versus-training tradeoff that mixed-use crawlers do, and says no well-established directive yet exists for expressing a training preference to an agent, though it points to an emerging standard called ai-prefs and says it will revisit the question as that matures. Cloudflare also previews its next target, AI Summaries, where it wants to let a site owner control how much of their content can appear in a summary, not only whether it appears at all, set once on Cloudflare rather than separately with each operator; the company's stated goal for that is early next year.
Key facts
- Cloudflare's new Disallow AI Training setting lets a mixed-use crawler, one that serves both search indexing and AI training, keep crawling a site for search while being blocked from training on it.
- Applebot, Bingbot, and Googlebot (run by Apple, Google, and Microsoft) qualify for Cloudflare's new "Accountable" designation; the separate, dedicated Training crawlers run by Amazon, Anthropic, Meta, and OpenAI are also categorized Accountable because blocking them never touches search.
- Cloudflare says less than 1% of its sites choose to block Search bots, versus 17% that enable some mechanism to block AI training.
- From September 15, Block and Block on pages with ads start applying to mixed-use crawlers too, and the old "Block AI Bots" toggle and Managed Robots.txt are deprecated in favor of granular Search, Training, and Agent controls and Bot Preference Sync.
- Microsoft's robots.txt-level no-training mechanism for Bing is targeted for early 2027 and is not live yet, so Disallow AI Training does not currently convey a no-training preference to Bingbot on its own.
Why it matters
Mixed-use crawlers used to force a single choice on a site owner: allow a crawler to train AI models on your content, or lose that crawler's contribution to your search visibility too, because refusing one meant refusing both. Cloudflare's new Disallow AI Training setting breaks that link for crawlers whose operators meet a bar the company calls Accountable, built from four commitments: a real opt-out mechanism for training, a real opt-out mechanism for AI summaries, visibility into which pages were used for training alongside search performance metrics, and assurance that opting out of training will not hurt a site's search ranking. Cloudflare says the bar itself came out of direct conversations with crawler operators since July, in which it reports that almost all of them agreed site owners should have control and transparency over how their content gets used. That turns AI crawl control from a blunt allow-or-block switch into a standing set of behavioral conditions an operator has to keep meeting to stay welcome.
Who it affects
Anyone running a site behind Cloudflare, since the Search, Training, and Agent controls, along with Disallow AI Training, are configured per domain there. The crawlers named directly are Applebot, Bingbot, and Googlebot, the three mixed-use crawlers Cloudflare currently calls Accountable, plus the separate, dedicated Training crawlers that Amazon, Anthropic, Meta, and OpenAI already run apart from their search crawlers. Existing Cloudflare customers who never touched the granular controls get moved automatically to new settings based on their old "Block AI Bots" choice, and anyone who had set Training to Block or Block on pages with ads is moved to Disallow AI Training, which Cloudflare says keeps the practical effect of that earlier choice. A domain onboarding after September 15 is steered toward one of two default presets depending on whether the site is ad-funded, though either preset can be changed at any time.
How to use it
The Training, Search, and Agent controls sit at the domain level in Cloudflare's dashboard, and Disallow AI Training is one of the choices available for Training specifically, alongside Allow, Block on pages with ads, and Block; it is not offered for Search or Agent. Turning it on makes Bot Preference Sync publish a Disallow: entry for training in the site's robots.txt; an Accountable mixed-use crawler keeps indexing the site for search under that setting, while every other training crawler, including the dedicated ones from Amazon, Anthropic, Meta, and OpenAI, is blocked outright. A site owner who wants Applebot, Bingbot, and Googlebot gone completely, from search too, now has to pick Block explicitly, since Block and Block on pages with ads apply to mixed-use crawlers as of September 15. Existing configurations do not need to be touched: Cloudflare migrates them automatically to the new system's equivalent setting. Opting a site out of Bing specifically still needs its own separate step for now, the NOARCHIVE meta tag or Cloudflare's Block URLs and Content Removal tools, because Bing's own robots.txt no-training support is not live yet.
How solid is it
This is Cloudflare's own announcement about its own product change, and the two headline figures in it, less than 1% of sites blocking Search versus 17% blocking some form of training, are self-reported from Cloudflare's own network rather than from an independent or audited source. The assurances that opting out of training will not affect search ranking are attributed directly to Apple, Google, and Microsoft themselves; the article reports what each company told Cloudflare, not an independent test of ranking impact. Part of what currently earns the Accountable label is a timetable rather than a shipped capability: Apple's URL-level inspection tool is described as in-progress and targeted for next year, and Microsoft's robots.txt-level no-training mechanism is targeted for early 2027, so two of the three mixed-use operators have not yet delivered every requirement Accountable status is meant to certify.
Risks and caveats
The Accountable label currently rewards commitments as well as finished capabilities, so a crawler operator can be described as Accountable before every one of the four requirements is actually live, as with Apple's pending inspection tool and Microsoft's pending robots.txt support. Cloudflare's own argument is that a robots.txt directive cannot by itself verify who is crawling or stop an operator that ignores it, which means the system still depends on Cloudflare correctly identifying and classifying crawler traffic to work as described. There is no Disallow option yet for Agent crawlers, the user-directed tools that fetch a page on a person's behalf, because Cloudflare says no established standard exists for expressing that preference to them; it points to an emerging standard, ai-prefs, that it says it will revisit. An ads-only Disallow option does not exist either, since Cloudflare says the set of ad-serving pages changes too fast to publish in robots.txt. And Disallow AI Training only ever covers training: the separate question of how much of a site's content can appear inside an AI-generated summary is still unresolved, with Cloudflare's own finer-grained control for that not due until early next year.
“The better outcome is operators that don't make you choose at all.”
— Cloudflare