The Tradeoff That Broke the Open Web
For years, site owners faced a binary choice: let mixed-use crawlers (Applebot, Googlebot, Bingbot) index your content for search — which also meant they could train models on it — or block them entirely and lose search discoverability.
That's a terrible deal. Less than 1% of Cloudflare sites block Search bots, but 17% actively try to block training. The gap tells you everything: publishers want to be found, they just don't want to be consumed.
The problem is structural. A robots.txt directive is a polite request. It can't identify who's crawling, classify why they're crawling, or stop a crawler that simply ignores it. Cloudflare's answer is to solve it at the network layer — publish the preference, identify the crawler, classify the intent, block the ones that ignore, and report on Radar what each operator actually does.
What Actually Changed on September 15
Cloudflare deprecated the old "Block AI Bots" toggle and replaced it with three granular, domain-level controls:
- Search — crawling to build a search index
- Training — crawling to train or fine-tune a model
- Agent — user-directed agents fetching pages on behalf of a human
The critical new setting is Disallow AI Training. It publishes a Disallow: directive through Bot Preference Sync into your robots.txt, keeps Accountable mixed-use crawlers allowed for search, and blocks every training-only crawler (Amazon, Anthropic, Meta, OpenAI).
Migration Table (Legacy → New)
| Legacy "Block AI" | Search | Training | Agent |
|---|---|---|---|
| Disabled | Allow | Allow | Allow |
| Block | Allow | Disallow AI Training | Block on pages with ads |
| Block on pages with ads | Allow | Disallow AI Training | Block on pages with ads |
If you previously set Training to Block or Block on pages with ads under the granular controls, both migrate to Disallow AI Training — preserving the practical effect without nuking search.
Recommended Presets for New Domains
| Setting | No ad monetization | Ad-monetized |
|---|---|---|
| Preference Sync | Enabled | Enabled |
| Search | Allow | Allow |
| Training | Allow | Disallow AI Training |
| Agent | Allow | Block on pages with ads |
Ad-supported sites get the stricter preset because ad revenue requires a human to actually land on the page. Training replaces that visit with a synthesized answer; agents fetch the page with nobody there to see the ads.

Operator-by-Operator: What Actually Works Today
Not every "Accountable" crawler is equal. The designation recognizes both shipped capabilities and time-bound commitments — which means some of it is vaporware until the delivery date lands.
Applebot
Opt out of training by adding a Disallow rule for Applebot-Extended in robots.txt. AI Summary preferences are expressible via the nosnippet directive in page HTML. Paywalled content can be labeled to exclude it from generative output.
Gap: No URL-level inspection tool yet. Apple has shared details of an in-progress solution for next year. They've stated disallowing training does not impact search ranking.
# robots.txt — Applebot training opt-out
User-agent: Applebot-Extended
Disallow: /
Googlebot
Opt out via Disallow for Google-Extended in robots.txt, plus a toggle in Search Console to exclude content from generative search results. Google provides metrics for both search and AI summary results.
Gap: URL-level transparency tools for Google-Extended are "expected in the weeks to come" — not shipped.
# robots.txt — Googlebot training opt-out (search unaffected)
User-agent: Google-Extended
Disallow: /
Bingbot
Currently only supports training opt-out via the NOARCHIVE meta tag or Microsoft's Block URLs / Content Removal tool. A robots.txt-level no-training preference is targeted for early 2027.
Critical caveat: Until that ships, selecting Disallow AI Training on Cloudflare will not automatically convey a no-training preference to Bing through robots.txt. This is the same practical behavior as the old Training Block setting.
<!-- Per-page training opt-out for Bingbot -->
<meta name="robots" content="NOARCHIVE">
Training-Only Crawlers (Amazon, Anthropic, Meta, OpenAI)
These operators separate their Search and Training crawlers, so Cloudflare can block the training crawler without touching search. No tradeoff, no negotiation.

Limitations and Things Nobody's Saying Out Loud
1. "Accountable" is a marketing designation, not a standard
Cloudflare invented the term. There's no IETF RFC for it. The requirements are self-defined and the enforcement is "we'll track it on Radar." That's better than nothing, but it's not a compliance regime.
2. Bing's 2027 commitment is a long time
If you're a publisher relying on Disallow AI Training today, Bingbot is still training on your content until Microsoft ships robots.txt support. The NOARCHIVE meta tag is per-page and easy to miss.
3. AI Summaries are the unsolved half
The current controls are binary: allow or block summaries. Cloudflare acknowledges this is "a blunt instrument" and says the next focus is letting you control how much of your content appears in a summary. That capability doesn't exist yet — it's a stated goal for early next year.
4. The conversion data cuts both ways
Cloudflare's own numbers: over half of consumers read summaries in Search, and those consumers are 40%+ more likely to end their search after reading one. But AI-referred traffic converts at 3×–5× the rate of traditional search referrals. Fewer visits, higher intent. Whether that's good or bad depends entirely on your business model — a publisher funded by ad impressions and a DTC retailer will draw opposite conclusions from the same data.
5. Agents have no directive yet
There's no established standard for expressing Disallow preferences to agents. Cloudflare is waiting on ai-prefs and similar proposals to mature. Until then, agent control is limited to "Block on pages with ads" or nothing.
Where to Go Next
- Audit your current settings at the domain (zone) level under Security Settings. If you never touched the granular controls, you've been migrated based on your legacy Block AI Bots state — verify it matches your intent.
- If you want mixed-use crawlers gone entirely, you now have to explicitly select Block. It stops Applebot, Bingbot, and Googlebot — search included.
- Watch the
ai-prefsspecification at the IETF. That's where the real interoperability battle is happening. - If you're building crawler-adjacent infrastructure, read the AI SDK 7 deep dive to understand how agent frameworks are evolving alongside these controls.
- For multilingual embedding and RAG pipelines that ingest web content, the IBM Granite Embedding Multilingual R2 breakdown covers the model-side implications of training data restrictions.

The Bottom Line
Cloudflare's Disallow AI Training is the first serious attempt to break the false binary between "indexed" and "trained on." The technical mechanism — network-level classification + robots.txt preference sync — is the right shape. The gaps are real: Bing's 2027 timeline, the missing URL-level transparency from Apple and Google, and the entirely unsolved AI Summaries problem.
If you run a content site, the immediate action is simple: verify your migrated settings, and if you're ad-monetized, make sure Disallow AI Training is on. If you're building on top of crawler data, start designing for a world where training access is opt-in, not default.