The Tradeoff That Broke the Open Web

For years, site owners faced a binary choice: let mixed-use crawlers (Applebot, Googlebot, Bingbot) index your content for search — which also meant they could train models on it — or block them entirely and lose search discoverability.

That's a terrible deal. Less than 1% of Cloudflare sites block Search bots, but 17% actively try to block training. The gap tells you everything: publishers want to be found, they just don't want to be consumed.

The problem is structural. A robots.txt directive is a polite request. It can't identify who's crawling, classify why they're crawling, or stop a crawler that simply ignores it. Cloudflare's answer is to solve it at the network layer — publish the preference, identify the crawler, classify the intent, block the ones that ignore, and report on Radar what each operator actually does.

근거자료: Cloudflare Blog — Accountable Mixed-Use AI Crawlers

What Actually Changed on September 15

Cloudflare deprecated the old "Block AI Bots" toggle and replaced it with three granular, domain-level controls:

  • Search — crawling to build a search index
  • Training — crawling to train or fine-tune a model
  • Agent — user-directed agents fetching pages on behalf of a human

The critical new setting is Disallow AI Training. It publishes a Disallow: directive through Bot Preference Sync into your robots.txt, keeps Accountable mixed-use crawlers allowed for search, and blocks every training-only crawler (Amazon, Anthropic, Meta, OpenAI).

Migration Table (Legacy → New)

Legacy "Block AI"SearchTrainingAgent
DisabledAllowAllowAllow
BlockAllowDisallow AI TrainingBlock on pages with ads
Block on pages with adsAllowDisallow AI TrainingBlock on pages with ads

If you previously set Training to Block or Block on pages with ads under the granular controls, both migrate to Disallow AI Training — preserving the practical effect without nuking search.

Recommended Presets for New Domains

SettingNo ad monetizationAd-monetized
Preference SyncEnabledEnabled
SearchAllowAllow
TrainingAllowDisallow AI Training
AgentAllowBlock on pages with ads

Ad-supported sites get the stricter preset because ad revenue requires a human to actually land on the page. Training replaces that visit with a synthesized answer; agents fetch the page with nobody there to see the ads.

Network diagram showing mixed-use AI crawlers split between search indexing and model training traffic Algorithm Concept Visual

Operator-by-Operator: What Actually Works Today

Not every "Accountable" crawler is equal. The designation recognizes both shipped capabilities and time-bound commitments — which means some of it is vaporware until the delivery date lands.

Applebot

Opt out of training by adding a Disallow rule for Applebot-Extended in robots.txt. AI Summary preferences are expressible via the nosnippet directive in page HTML. Paywalled content can be labeled to exclude it from generative output.

Gap: No URL-level inspection tool yet. Apple has shared details of an in-progress solution for next year. They've stated disallowing training does not impact search ranking.

# robots.txt — Applebot training opt-out
User-agent: Applebot-Extended
Disallow: /

Googlebot

Opt out via Disallow for Google-Extended in robots.txt, plus a toggle in Search Console to exclude content from generative search results. Google provides metrics for both search and AI summary results.

Gap: URL-level transparency tools for Google-Extended are "expected in the weeks to come" — not shipped.

# robots.txt — Googlebot training opt-out (search unaffected)
User-agent: Google-Extended
Disallow: /

Bingbot

Currently only supports training opt-out via the NOARCHIVE meta tag or Microsoft's Block URLs / Content Removal tool. A robots.txt-level no-training preference is targeted for early 2027.

Critical caveat: Until that ships, selecting Disallow AI Training on Cloudflare will not automatically convey a no-training preference to Bing through robots.txt. This is the same practical behavior as the old Training Block setting.

<!-- Per-page training opt-out for Bingbot -->
<meta name="robots" content="NOARCHIVE">

Training-Only Crawlers (Amazon, Anthropic, Meta, OpenAI)

These operators separate their Search and Training crawlers, so Cloudflare can block the training crawler without touching search. No tradeoff, no negotiation.

Cloudflare dashboard security settings panel displaying Disallow AI Training toggle for domain-level bot management System Abstract Visual

Limitations and Things Nobody's Saying Out Loud

1. "Accountable" is a marketing designation, not a standard

Cloudflare invented the term. There's no IETF RFC for it. The requirements are self-defined and the enforcement is "we'll track it on Radar." That's better than nothing, but it's not a compliance regime.

2. Bing's 2027 commitment is a long time

If you're a publisher relying on Disallow AI Training today, Bingbot is still training on your content until Microsoft ships robots.txt support. The NOARCHIVE meta tag is per-page and easy to miss.

3. AI Summaries are the unsolved half

The current controls are binary: allow or block summaries. Cloudflare acknowledges this is "a blunt instrument" and says the next focus is letting you control how much of your content appears in a summary. That capability doesn't exist yet — it's a stated goal for early next year.

4. The conversion data cuts both ways

Cloudflare's own numbers: over half of consumers read summaries in Search, and those consumers are 40%+ more likely to end their search after reading one. But AI-referred traffic converts at 3×–5× the rate of traditional search referrals. Fewer visits, higher intent. Whether that's good or bad depends entirely on your business model — a publisher funded by ad impressions and a DTC retailer will draw opposite conclusions from the same data.

5. Agents have no directive yet

There's no established standard for expressing Disallow preferences to agents. Cloudflare is waiting on ai-prefs and similar proposals to mature. Until then, agent control is limited to "Block on pages with ads" or nothing.

Where to Go Next

  • Audit your current settings at the domain (zone) level under Security Settings. If you never touched the granular controls, you've been migrated based on your legacy Block AI Bots state — verify it matches your intent.
  • If you want mixed-use crawlers gone entirely, you now have to explicitly select Block. It stops Applebot, Bingbot, and Googlebot — search included.
  • Watch the ai-prefs specification at the IETF. That's where the real interoperability battle is happening.
  • If you're building crawler-adjacent infrastructure, read the AI SDK 7 deep dive to understand how agent frameworks are evolving alongside these controls.
  • For multilingual embedding and RAG pipelines that ingest web content, the IBM Granite Embedding Multilingual R2 breakdown covers the model-side implications of training data restrictions.

Server infrastructure routing crawler requests through robots.txt preference sync for AI training opt-out Coding Session Visual

The Bottom Line

Cloudflare's Disallow AI Training is the first serious attempt to break the false binary between "indexed" and "trained on." The technical mechanism — network-level classification + robots.txt preference sync — is the right shape. The gaps are real: Bing's 2027 timeline, the missing URL-level transparency from Apple and Google, and the entirely unsolved AI Summaries problem.

If you run a content site, the immediate action is simple: verify your migrated settings, and if you're ad-monetized, make sure Disallow AI Training is on. If you're building on top of crawler data, start designing for a world where training access is opt-in, not default.

This content was drafted using AI tools based on reliable sources, and has been reviewed by our editorial team before publication. It is not intended to replace professional advice.