All news
aicybersecurityproductregulation

Cloudflare Sets New AI Bot Defaults for Sept 2026

28 Jul 2026

Cloudflare is overhauling how it manages AI crawler traffic across its network, one year after debuting its original "Block AI Bots" toggle and Pay-Per-Crawl marketplace. The update introduces a more granular classification system and sets new default blocking rules that will take effect for new domains starting September 15, 2026.

What's changing

Cloudflare is moving from a single blanket "block AI bots" switch to a three-category framework based on crawler use case:

  • Search — crawlers indexing content for search results
  • Agent — crawlers acting on behalf of AI agents
  • Training — crawlers collecting data to train AI models

Starting September 15, 2026, any new domain onboarding to Cloudflare will have Training and Agent crawlers blocked by default on pages that display ads. Search crawlers will remain allowed by default on those same ad-displaying pages.

Alongside this, Cloudflare is launching BotBase, a new feature inside Enterprise Bot Management that tracks known bots — including verified bots and agents. Cloudflare says BotBase will later be expanded into a broader control center for managing known automated traffic on websites, though no technical details on that expansion have been shared yet.

Importantly, these AI traffic management options — the new default classifications and related controls — are being made available to all Cloudflare network customers, including those on the Free tier, not just enterprise accounts.

The complication: multi-purpose crawlers

One wrinkle in the new system: some well-known crawlers don't fit neatly into a single category. Cloudflare notes that crawlers like Googlebot, Applebot, and BingBot combine both Search and Training functions. Site owners who choose to block Training crawlers under the new framework will end up blocking these multi-purpose bots too — potentially reducing search visibility as a side effect of trying to restrict AI training access.

Risks and open questions

Several practical risks remain unresolved:

  • Blocking Training by default could inadvertently catch multi-purpose crawlers, hurting search discoverability for some sites.
  • The transition period before September 2026 may produce inconsistent bot treatment as domains onboard under different rules.
  • Ad-display detection — the mechanism that determines whether default blocking applies — could misfire, either over-blocking or under-blocking AI crawlers.

Cloudflare hasn't detailed how existing (already onboarded) domains will be treated relative to new ones, how it defines or detects "ads" for this purpose, or how site owners can override the new defaults. There's also no adoption or usage data available yet from the first year of Block AI Bots or Pay-Per-Crawl, making it hard to gauge how much real-world impact these tools have had so far.

Why founders should care

For early-stage teams, these changes are likely to matter in a few concrete ways:

  • If your startup depends on AI-driven content indexing or discovery, you may need to check how your site is categorized under the new Search/Agent/Training framework — misclassification could plausibly affect visibility.
  • Founders running ad-supported sites should expect reduced AI training crawler access by default, which may narrow future data-licensing opportunities unless settings are adjusted.
  • Companies building AI agents that crawl external content should verify their bots are properly classified, since default rules could block them starting in September 2026.
  • Pay-Per-Crawl and BotBase could offer early-stage companies new ways to monetize or manage automated traffic, but without adoption data, it's uncertain how significant this opportunity will turn out to be in practice.

The bottom line

Cloudflare is pushing site owners toward more deliberate, granular control over AI crawler access — a shift that could reshape how startups think about both AI training exposure and search visibility. With the new defaults landing more than a year out, founders have time to review their crawler policies, but the lack of clarity around ad detection, override options, and treatment of existing domains means questions remain heading into the transition.

Sources