Cloudflare’s September Bot Defaults Could Quietly Cut AI Visibility
By Greg Nowak. Last updated 2026-08-26.
Cloudflare is giving website owners a clearer say over how automated systems use their content. Publishers have good reason to welcome that. The catch is that a rule meant to prevent AI training may also shut out a crawler that helps people find the site.
From September 15, 2026, new domains joining Cloudflare’s network will receive behavior-based defaults. On pages displaying advertising, Search crawlers will be allowed by default, while Agent and Training crawlers will be blocked. That looks tidy on a settings screen. Crawler identities are less tidy.
Cloudflare says a multipurpose crawler is governed by its most restrictive applicable rule. If the same crawler is classified for both Search and Training, blocking Training can block it altogether. Cloudflare names Googlebot, Applebot and BingBot as examples. A publisher could make a reasonable content-protection decision and inadvertently affect indexing, search referrals or newer forms of AI discovery.
This is a distribution policy, not just a bot setting
Bot controls used to sit comfortably with the infrastructure team: stop abusive automation, reduce unwanted traffic and protect server capacity. Cloudflare’s taxonomy puts a commercial decision inside that technical configuration.
The platform separates automated traffic into three main behaviors. Search collects or indexes content to answer later queries. Agent traffic acts for a person, often in real time. Training gathers material to train or fine-tune models. One crawler may have more than one classification, although Cloudflare is encouraging crawler operators to use separate identities for separate purposes.
That separation could let publishers remain discoverable without automatically permitting training. Until crawler operators adopt it consistently, “Block Training” is not always the narrow instruction it appears to be.
The people responsible for publishing, acquisition and licensing therefore need a say. The configuration decides which services can discover the content, what they may do with it and whether that access is expected to produce referrals or another form of value.
| Traffic behavior | Default on ad-supported pages for new domains | Business decision | Practical check |
|---|---|---|---|
| Search | Allowed | Which discovery services matter? | Confirm priority pages remain crawlable and indexable |
| Agent | Blocked | Could user-directed agents support a useful customer journey? | Test deliberate exceptions on the intended pages |
| Training | Blocked | Is training prohibited, permitted or covered by an agreement? | Check whether the rule also blocks a mixed-purpose search crawler |
| Mixed-purpose crawler | Most restrictive applicable rule | Does its discovery value justify a scoped exception? | Review requests, rule outcomes and referral changes |
Content protection and visibility can coexist
The concern about uncompensated use is real. TechCrunch reports that Cloudflare is expanding its commercial model from Pay Per Crawl towards Pay Per Use, under which participating publishers could be paid when their content creates value rather than only when it is fetched. The report also describes early arrangements involving Ceramic.ai and You.com. Automated access is becoming something publishers can govern and, in some cases, license.
Recent experimental research adds weight to the referral question. In a preregistered field experiment with 1,100 participants, researchers tested Google AI Overviews and AI Mode. Removing those AI features increased click-through rates to publishers. An AI Mode-only experience reduced publisher click-through and weakened reported user experience and trust. That does not predict the outcome for every website, but it is a sound reason to measure referrals instead of assuming AI exposure will replace lost visits.
A blanket block creates a different problem. TechRadar points to the risk of denying a mixed-purpose crawler such as Googlebot: conventional search discovery and page updates may suffer too. The attempt to preserve human traffic could make the content harder for humans to find.
The workable goal is controlled distribution. Allow access tied to a defined outcome, restrict uses the publisher has not accepted, and check what actually happens in the traffic data.
Audit the whole control path
A reliable review starts with the site’s effective behavior, not a screenshot of one Cloudflare toggle. Several layers can declare or enforce crawler policy, and they need to agree.
Start by listing the domains and page types in scope. The announced defaults apply to new domains and distinguish pages that display advertising. That makes advertising templates and deployment plans relevant. A staging hostname, regional domain or newly launched publication could otherwise inherit a policy that nobody responsible for acquisition has reviewed.
Next, record the Cloudflare choices for Search, Agent and Training traffic, along with any legacy “Block AI bots” configuration. Cloudflare allows customers to choose settings other than the new defaults. The point is not to accept or reject Cloudflare’s position wholesale; it is to make the organisation’s own choice explicit.
Then compare the enforcement rules with robots.txt and Content-Signal directives. Cloudflare’s managed robots.txt can express preferences such as allowing search, disallowing AI training, and limiting use to reference-level indexing, excerpting and linking. But robots.txt states a preference; it does not create the network-level block. A declared policy that contradicts the enforced rule will be difficult to explain and harder to operate.
Pay particular attention to multipurpose crawlers the business relies on. An allowed Search category does not guarantee access if that crawler also carries a blocked Training classification. Test representative URLs, including valuable editorial pages, newly published articles, advertising templates and paths covered by special rules.
Before changing anything, capture a baseline of crawler requests, conventional search referrals and identifiable AI referrals. Afterwards, monitor rule outcomes and discovery signals. The comparison will not prove causation by itself, but it can expose a sudden shift early enough to roll back the change or create a narrower exception.
Give the policy an owner and a review date
Crawler classifications will change, as will the AI products using them. Cloudflare’s taxonomy is behavior-based precisely because a simple “AI or not AI” label no longer describes many modern services. A one-off configuration review will age quickly.
The written policy should record the business objective, permitted behaviors, prohibited uses, exceptions, affected domains, implementation layers, retained evidence and the person accountable for review. It also needs clear reassessment triggers: a new domain, a changed advertising template, a crawler reclassification, a material referral shift or a new content-licensing agreement.
Cloudflare’s defaults may be a useful starting point, but they are still defaults. Greg can help turn them into a documented allow/block policy, test the crawlers that matter and put lightweight monitoring around the result. That gives publishers a practical way to protect their content without quietly giving up valuable search or AI visibility.
Related on GrN.dk
- Google’s AI Search Toggle Is a Publishing Decision, Not an SEO Setting
- Search Console Can See TikTok Now. Your Reporting Has to Catch Up
- Your AI workflow has logs. Can they explain one bad decision?
Need help with this kind of work?
Review your crawler policy Get in touch with Greg.