Growth & Strategy

On September 15 the Web Starts Hiding From AI, and It Takes Your Press Coverage With It

August 24, 2026

From September 15, Cloudflare blocks AI crawlers on ad-supported pages by default, thinning the independent coverage AI systems can read about your category.

On September 15 the Web Starts Hiding From AI, and It Takes Your Press Coverage With It
Credit:
powered by

Make State of Brand one of your go-to sources on Google

Google Icon
Add State of Brand on Google

Cloudflare put out a post on July 1 that most marketing teams filed under infrastructure and never took seriously.

Halfway down it says that from September 15, multi-purpose crawlers get allowed or blocked according to all of their behaviors, with the most restrictive rule winning. Cloudflare names the consequence itself. Googlebot, Applebot and BingBot will be blocked for any customer who has chosen to block Training, whether through the new controls or through the older Block AI bots service.

That older button is the one a lot of companies pressed in 2024 and 2025, during the year everyone got nervous about their content ending up in a training set. It was one click, it sat in the dashboard beside things that made obvious sense, and pressing it was not a decision about Google Search indexing, because at the time it wasn't one.

Cloudflare sits in front of more than 20% of web domains. The number of sites that pressed that button and forgot about it will not be small.

What Cloudflare actually changed

The company has stopped sorting bots into AI and not-AI, on the grounds that the line keeps moving. Google Search now answers questions on the results page instead of sorting links, and Cloudflare's post points at that as the reason the old category stopped meaning much.

The replacement taxonomy asks what a bot does once it arrives. Search collects and indexes content so a system can answer questions about it later, and Cloudflare argues site owners should get referral traffic or equivalent compensation for it. Agent covers automation acting in real time for a person, including chat fetch bots and the agents driving browsers, usually with a human waiting at the other end. Training covers crawlers absorbing content permanently into a model. Cloudflare tracks eight more categories beyond those three, among them Transact, Data Collection, SEO and Ads Verification, and it is pushing operators to split their crawlers by purpose so site owners can tell why something showed up at the door.

The defaults change on September 15. For all new domains onboarding to Cloudflare, Training and Agent get blocked on pages that display ads, and Search stays allowed. The logic is that an ad signals the owner meant a person to land there, so those pages treat human attention as the point. TechCrunch reported the new defaults reaching further than Cloudflare's own wording suggests, to new sites from existing customers and to all existing free accounts.

Opting out takes a few clicks in Security settings any time before September 15, and confirms no changes to Training crawlers that also crawl for Search.

The part that matters for B2B brands

Run the ad-page rule against your own category and see where it lands.

Your company blog almost certainly carries no ads. Neither does your pricing page, your documentation or your customer stories. None of it gets blocked by the new default, so your marketing keeps flowing into training corpora and agent retrieval exactly as it does now.

Ads live on trade publications, industry blogs, review roundups, analyst commentary, the newsletters your category actually reads. Those are the pages that mention you without your involvement, and by default they are the ones closing to Training and Agent crawlers.

The mix shifts accordingly. Material an AI system can reach about your category gets heavier on vendor-published content and lighter on independent coverage, starting on one specific Tuesday, because of a setting almost nobody will touch.

This publication faces the same decision, and pretending otherwise would be silly.

Content use, and a new line in robots.txt

Cloudflare is also testing a signal covering what a crawler may keep and reshare after reading a page. Three levels. Immediate means interact and store nothing. Reference, the default, means index, excerpt and link back. Full means summarize and reproduce.

The signal extends Content Signals and lives in robots.txt. Sites already using Cloudflare's managed robots.txt now get use=reference appended to the existing search and training preferences. Cloudflare says it will track content use for every bot in its directory, and that bots caught abusing the signals lose Verified status. Bots that reproduce in full cannot be Verified at all today.

Verified has been redefined too. It no longer means default allowed. Non-verified bots stay blocked, and Verified now makes a bot allowable inside whichever category the site owner has permitted.

Where this could be wrong

Cloudflare's reasoning holds up better than the headline suggests.

The company frames the old arrangement as a bargain that broke. Crawlers took content and sent referrals back for thirty years, and AI kept only the first half of that. Its post describes the position small sites ended up in as a choice between appearing in search while accepting training, or losing discoverability, which hands an advantage to incumbents running one crawler for both jobs. Forcing operators to separate crawlers by purpose answers that reasonably well.

This is one company setting defaults, not legislation. Anyone who wants the old behavior keeps it with a few clicks, and Cloudflare says it will keep warning customers as the date approaches.

Against all of that, defaults decide outcomes for the large majority who never open the settings, and this one reshapes what AI systems can read about entire categories.

Three weeks

Open your Cloudflare dashboard and find out what your site is doing today, since the answer may predate everyone currently on your team.

The specific thing to look for is whether Block AI bots was ever enabled on your zone. If it was, Googlebot, Applebot and BingBot are in scope on September 15 unless somebody opts out.

Agent traffic deserves a deliberate decision instead of an inherited setting. Chat fetch bots and browser-driving agents show up with a buyer attached, and blocking them removes you from the sessions where evaluation happens.

Then ask the publications that cover your category what they plan to do, because their settings shape what a model can say about you and you have no vote in it.

Outlever Logo

If this caught your attention, that’s not accidental.


Decoration line

The best editorial systems don’t happen by accident. Outlever builds them.

Partial view of green concentric circles with a solid green dot on the outermost circle on a light background.Concentric green circles with a single solid green dot on a dashed circle on a light background.Minimalist design with faint curved lines and scattered small green dots on a white background.

Come back for the reason it lands.


Subscribe for the kind of thinking that makes people stop, read and come back.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.