Cloudflare’s September Bot Defaults Could Quietly Cut AI Visibility

Illustrated infographic summarizing: Cloudflare’s September Bot Defaults Could Quietly Cut AI Visibility

By Greg Nowak. Last updated 2026-08-26.

Cloudflare is giving website owners a clearer say over how automated systems use their content. Publishers have good reason to welcome that. The catch is that a rule meant to prevent AI training may also shut out a crawler that helps people find the site.

From September 15, 2026, new domains joining Cloudflare’s network will receive behavior-based defaults. On pages displaying advertising, Search crawlers will be allowed by default, while Agent and Training crawlers will be blocked. That looks tidy on a settings screen. Crawler identities are less tidy.

Cloudflare says a multipurpose crawler is governed by its most restrictive applicable rule. If the same crawler is classified for both Search and Training, blocking Training can block it altogether. Cloudflare names Googlebot, Applebot and BingBot as examples. A publisher could make a reasonable content-protection decision and inadvertently affect indexing, search referrals or newer forms of AI discovery.

This is a distribution policy, not just a bot setting

Bot controls used to sit comfortably with the infrastructure team: stop abusive automation, reduce unwanted traffic and protect server capacity. Cloudflare’s taxonomy puts a commercial decision inside that technical configuration.

The platform separates automated traffic into three main behaviors. Search collects or indexes content to answer later queries. Agent traffic acts for a person, often in real time. Training gathers material to train or fine-tune models. One crawler may have more than one classification, although Cloudflare is encouraging crawler operators to use separate identities for separate purposes.

That separation could let publishers remain discoverable without automatically permitting training. Until crawler operators adopt it consistently, “Block Training” is not always the narrow instruction it appears to be.

The people responsible for publishing, acquisition and licensing therefore need a say. The configuration decides which services can discover the content, what they may do with it and whether that access is expected to produce referrals or another form of value.

Traffic behavior Default on ad-supported pages for new domains Business decision Practical check
Search Allowed Which discovery services matter? Confirm priority pages remain crawlable and indexable
Agent Blocked Could user-directed agents support a useful customer journey? Test deliberate exceptions on the intended pages
Training Blocked Is training prohibited, permitted or covered by an agreement? Check whether the rule also blocks a mixed-purpose search crawler
Mixed-purpose crawler Most restrictive applicable rule Does its discovery value justify a scoped exception? Review requests, rule outcomes and referral changes
A working matrix for turning Cloudflare’s crawler defaults into a deliberate publishing policy.

Content protection and visibility can coexist

The concern about uncompensated use is real. TechCrunch reports that Cloudflare is expanding its commercial model from Pay Per Crawl towards Pay Per Use, under which participating publishers could be paid when their content creates value rather than only when it is fetched. The report also describes early arrangements involving Ceramic.ai and You.com. Automated access is becoming something publishers can govern and, in some cases, license.

Recent experimental research adds weight to the referral question. In a preregistered field experiment with 1,100 participants, researchers tested Google AI Overviews and AI Mode. Removing those AI features increased click-through rates to publishers. An AI Mode-only experience reduced publisher click-through and weakened reported user experience and trust. That does not predict the outcome for every website, but it is a sound reason to measure referrals instead of assuming AI exposure will replace lost visits.

A blanket block creates a different problem. TechRadar points to the risk of denying a mixed-purpose crawler such as Googlebot: conventional search discovery and page updates may suffer too. The attempt to preserve human traffic could make the content harder for humans to find.

The workable goal is controlled distribution. Allow access tied to a defined outcome, restrict uses the publisher has not accepted, and check what actually happens in the traffic data.

Audit the whole control path

A reliable review starts with the site’s effective behavior, not a screenshot of one Cloudflare toggle. Several layers can declare or enforce crawler policy, and they need to agree.

Start by listing the domains and page types in scope. The announced defaults apply to new domains and distinguish pages that display advertising. That makes advertising templates and deployment plans relevant. A staging hostname, regional domain or newly launched publication could otherwise inherit a policy that nobody responsible for acquisition has reviewed.

Next, record the Cloudflare choices for Search, Agent and Training traffic, along with any legacy “Block AI bots” configuration. Cloudflare allows customers to choose settings other than the new defaults. The point is not to accept or reject Cloudflare’s position wholesale; it is to make the organisation’s own choice explicit.

Then compare the enforcement rules with robots.txt and Content-Signal directives. Cloudflare’s managed robots.txt can express preferences such as allowing search, disallowing AI training, and limiting use to reference-level indexing, excerpting and linking. But robots.txt states a preference; it does not create the network-level block. A declared policy that contradicts the enforced rule will be difficult to explain and harder to operate.

Pay particular attention to multipurpose crawlers the business relies on. An allowed Search category does not guarantee access if that crawler also carries a blocked Training classification. Test representative URLs, including valuable editorial pages, newly published articles, advertising templates and paths covered by special rules.

Before changing anything, capture a baseline of crawler requests, conventional search referrals and identifiable AI referrals. Afterwards, monitor rule outcomes and discovery signals. The comparison will not prove causation by itself, but it can expose a sudden shift early enough to roll back the change or create a narrower exception.

Give the policy an owner and a review date

Crawler classifications will change, as will the AI products using them. Cloudflare’s taxonomy is behavior-based precisely because a simple “AI or not AI” label no longer describes many modern services. A one-off configuration review will age quickly.

The written policy should record the business objective, permitted behaviors, prohibited uses, exceptions, affected domains, implementation layers, retained evidence and the person accountable for review. It also needs clear reassessment triggers: a new domain, a changed advertising template, a crawler reclassification, a material referral shift or a new content-licensing agreement.

Cloudflare’s defaults may be a useful starting point, but they are still defaults. Greg can help turn them into a documented allow/block policy, test the crawlers that matter and put lightweight monitoring around the result. That gives publishers a practical way to protect their content without quietly giving up valuable search or AI visibility.

Related on GrN.dk

Need help with this kind of work?

Review your crawler policy Get in touch with Greg.

Sources

Latest articles

When an OpenAI request stalls, customers need an accurate status. Set sensible retry limits, preserve submissions, and make unresolved work visible.

I learned server operations by breaking my own servers. I want someone who stands next to me while I do it, then does it themselves the week after.

I am good at building and bad at calling. Here is who I want next to me, what is easiest to sell, and how we split it.

An AI assistant can prepare a refund, but a person should approve the exact payment and amount. Here is how to make that approval hold up through execution and retries.

AI can pull together onboarding tasks before a new hire’s first day. See how the manager approves specific access and how outstanding tasks are followed through.

An internal AI assistant can cite an obsolete handbook with confidence. Here is how to manage document ownership, updates, deletions, access and answer review.

Cloudflare Free provides useful website protection, but its rate limiting and bot controls have limits. Here is how to assess them for a WordPress site.

An AI assistant can answer questions and guide customers to a booking. Here are practical boundaries for prices, delivery times, personal data, and contact with a staff member.

Google and Bing now offer first-party AI search visibility reports. Here’s how to build a useful baseline without inventing a misleading GEO score.

AI crawlers can copy a familiar name. Here’s how to verify signed agents at the edge while keeping legitimate automated traffic moving.