Cloudflare's AI block now stops everyone but Google, Apple and Microsoft

The week Cloudflare set for converting its customers' AI-training blocks into robots.txt preferences has run. Google, Apple and Microsoft are now merely asked to honor those preferences; everyone else's training crawlers still get stopped at the edge.

Vincent JiangVincent Jiang · 3 min read
Share
Matthew Prince, co-founder and chief executive of Cloudflare, speaking at TechCrunch Disrupt
1 / 7Slide 1 of 7
Cloudflare co-founder and CEO Matthew Prince. The company's week-long migration has turned customers' AI-training blocks into robots.txt preferences.

Cloudflare's customer email of 16 September 2026 said legacy Block AI Bots settings would migrate automatically over the following week, and that the switch would vanish from the dashboard once the migration was complete 1. That week has run. What publishers clicked as a block is now, for the three companies that matter most, a polite note in a text file.

What a block became

The mapping is fixed. A legacy toggle lands on Search Allow, Training Disallow AI Training, and Agent blocked on pages with ads 12. Disallow AI Training writes a no-training directive into robots.txt while Googlebot, Applebot and Bingbot keep fetching, because all three companies run one crawler for both search indexing and model training 12. The gap it addresses is real: 17% of sites on Cloudflare's network block training in some form, fewer than 1% block search bots, and the network sits in front of more than a fifth of the web 13.

Blocking training is common on Cloudflare's network; blocking search is rare

05101520Sites blocking AI training17%Sites blocking search crawlers<1%Web behind Cloudflare's network>20%
Data
Value
Sites blocking AI training17%
Sites blocking search crawlers<1%
Web behind Cloudflare's network>20%
Share of customer sites on Cloudflare's network that block AI training in some form (17%), and the fewer than 1% that block search crawlers; the network itself sits in front of more than a fifth of the web, so one dashboard setting scales accordingly. Sources: PPC Land, 27 September 2026; Best Media Info, 16 September 2026.1,3

Enforcement now has two tiers

The setting still stops someone. Training crawlers run by OpenAI, Anthropic, Meta and Amazon are blocked outright at Cloudflare's edge 4. Google, Apple and Microsoft keep fetching under Cloudflare's Accountable label, and what happens to the page afterwards rests on the operator honoring the file 14. Microsoft cannot even read the new rule until its robots.txt support arrives, targeted for early 2027 12.

Accountable is Cloudflare's own designation, not an outside certification 4. The label was granted on capabilities each operator already has, paired with time-bound commitments for the rest 12. And in August the fourth condition required an operator to "show publicly" that disallowing training does not hurt search results; the September version asks for assurance 51.

The manual still promises the opposite

Cloudflare's own documentation, last updated 1 July 2026, still says mixed-purpose crawlers "will also be blocked by all configurations to block AI training, including the legacy Block AI bots option" 6. That page lists three options for each control: Allow, Block on pages with ads, and Block 6. The shipped product carries a fourth, Disallow AI Training, which the manual never mentions 1. The behavior that shipped is the inverse of the page Cloudflare still publishes.

The only way back to a hard block is Block, which now cuts a site off from Google, Apple and Microsoft search altogether 12. Blocking Bingbot takes a site out of Bing, Yahoo and Copilot at once 1.

The receipt arrives in weeks, not now

URL-level reporting, the mechanism meant to reveal which disallowed pages were used anyway, is weeks away at Google and due next year at Apple 124. Until it lands, the ledger reads one way: publishers keep their search traffic, the giants keep the corpus, Cloudflare keeps both sides as customers, and the only crawlers still stopped by force belong to the giants' AI rivals 4. Summaries are next, with central controls targeted for early 2027 13.

How this brief was made

01Gathered & sourced272 channels · 1,762 articles▾

Agents swept 272 channels and ingested 1,762 articles, then de-duplicated and ranked them for signal.

02Verified & cross-validated6 claims · 22 data feeds▾
03Reviewed & edited1 human editor▾

One editor read the draft against the evidence, tuned the framing, and signed off before it shipped.

Become a contributor

Reporting on the business of AI and want it read? We take pitches from outside contributors who bring primary sources and a number worth arguing about.

Share

Deepdive

AI-generated from this story and its cited sources. Not investment advice.

Reader comments

0 comments

    Sign up

    Get your curated digest

    After email confirmation, you will receive a daily digest of the most relevant news that matter to your portfolio