CMOtech Asia - Technology news for CMOs & marketing decision-makers
Asia
Cloudflare lets sites block AI training but keep search

Cloudflare lets sites block AI training but keep search

Wed, 16th Sep 2026 (Today)
Sean Mitchell
SEAN MITCHELL Publisher

Cloudflare has launched a new setting called Disallow AI Training for website owners. It lets sites remain indexed for search while preventing their content from being used to train AI models.

The change addresses a growing problem caused by mixed-use crawlers, which collect material for both search indexing and AI training through a single bot. Those crawlers now account for 36.6% of all verified crawler traffic on Cloudflare's network, making them the largest category it tracks.

For website owners, the issue has become a trade-off between visibility and control. Cloudflare says fewer than 1% of websites block search crawlers, while 17% already restrict AI training. Until now, sites trying to stop AI training by mixed-use crawlers often risked losing visibility in search results as well.

The new setting is part of a broader update to Cloudflare's controls for automated access to websites. It replaces the earlier Block AI Bots switch with three separate controls covering search, AI training, and AI agents.

Cloudflare is also introducing recommended settings for new websites based on business model. Sites that carry advertising will, by default, have search crawling enabled, AI training disallowed, and AI agents blocked on ad-carrying pages. Other sites will have search, AI training, and AI agents allowed.

It is also replacing its Managed Robots.txt feature with a system called Bot Preference Sync. The tool lets website owners set crawling preferences once and apply them automatically across supported crawlers.

Accountable label

As part of the rollout, Cloudflare has created an Accountable designation for AI crawling. Apple, Google, and Microsoft are the first companies to receive the label for mixed-use crawlers after meeting Cloudflare's criteria or providing timelines to do so.

To qualify, operators must meet four conditions. They must give site owners a clear way to opt out of AI training through robots.txt or a comparable standard, allow opt-outs from AI-generated search summaries, provide URL-level visibility into how content is used for search and training, and publicly confirm that opting out of training will not affect a site's ranking in traditional search.

Other AI companies use separate crawlers for search and training. In those cases, Cloudflare can block training crawlers without disrupting search, and it also classifies the relevant crawlers as Accountable.

The move reflects a broader tension between publishers, online businesses, and AI developers over how web content is collected and reused. Search engines have long depended on crawling websites in return for directing users back to them. AI systems have complicated that arrangement by ingesting material for model training and, increasingly, for generated answers and summaries that can reduce referral traffic.

Cloudflare's figures suggest many website owners distinguish between being discoverable in search and contributing material to AI models. The rise of mixed-use crawlers has made that distinction harder to enforce because a single bot may serve both functions.

Matthew Prince, Chief Executive Officer of Cloudflare, said the new settings are intended to preserve that distinction for site owners. "This is how we make the Internet better: preserving the openness that makes search valuable while giving the people and businesses behind the web meaningful control over how their work is used," Prince said. "We look forward to continued collaborative engagement with companies like Apple, Google, and Microsoft as we work together to build a healthy ecosystem."

Next battleground

Cloudflare is also turning its attention to AI summaries, another area where web publishers have raised concerns. Operators it designates as Accountable must give site owners a way to opt out of AI summaries, the generated overviews that can appear in search experiences.

At present, that choice is handled on an operator-by-operator basis. Cloudflare wants site owners to be able to control how much of their content is included from a single place rather than managing settings separately across multiple providers.

Its current controls are designed to work alongside emerging industry standards, including ai-prefs, a specification being developed through the Internet Engineering Task Force. That work aims to create a more standardised and portable way for websites to express preferences over AI access.

Cloudflare's changes underline how infrastructure providers are taking a more active role in disputes over web content and AI use. By placing default settings, crawler classifications, and synchronised controls between site owners and bot operators, it is trying to shape the practical rules of access at the network level.

Its data shows why that matters: mixed-use crawlers are now the single biggest class of verified crawler traffic on Cloudflare's network, while a significant minority of websites already want to stop AI training without giving up search visibility.