On September 15, 2026 Cloudflare, which sits in front of roughly one in five websites, started treating AI crawlers as three different jobs and blocking some of them by default. If your website runs through Cloudflare, your settings may have changed without you touching them. If it doesn't, the same three-way split is becoming the way every AI company reads the web. Either way, your site is now expected to say what AI may do with it.
This guide is for business owners, not engineers. It covers what changed, who it affects, and the four decisions it asks you to make.
The three kinds of AI crawler
A crawler is a program that reads web pages. Until this year, most AI companies used one crawler for everything. Cloudflare now separates them by what they do with your pages:
| Kind | What it does | Examples | Cloudflare's default since Sept 15 |
|---|---|---|---|
| Search | Builds search results and sends people to your site with a link | Googlebot, Bingbot, OAI-SearchBot (ChatGPT search), PerplexityBot | Allowed |
| Agent | Opens a page because a specific person asked an AI assistant to | ChatGPT-User, Claude-User, Perplexity-User | Allowed, except on pages that show ads |
| Training | Copies content in bulk to train AI models. Nothing comes back to your site | GPTBot, ClaudeBot, Google-Extended, CCBot | Blocked on ad-supported sites, allowed elsewhere |
A simple way to hold it: search brings visitors, agents bring answers, training brings nothing back.
The new defaults apply to new Cloudflare customers, new sites added by existing customers, and every site on Cloudflare's free plan. Paid customers keep whatever they had unless they change it.
What else came with it
Three things, according to Cloudflare's announcement:
- Mixed-use crawlers are blocked by default on pages with ads. A mixed-use crawler collects for search and for training at the same time. Cloudflare says these were the single largest category of verified crawler traffic on its network. AI companies were given until September 15 to split them.
- Disallow AI Training is a one-click setting that refuses training while keeping the site fully in search results.
- Accountable is a label for AI companies that honour robots.txt opt-outs, let sites opt out of AI summaries, show which pages they fetched, and confirm publicly that opting out of training does not hurt search ranking. Apple, Google and Microsoft were named first.
Cloudflare also said its Pay Per Crawl marketplace is becoming Pay Per Use, which charges AI companies when content creates value rather than each time it is fetched. That is aimed at publishers. A dentist, a home care agency or a law firm does not need to act on it.
Will you disappear from Google or ChatGPT?
Not from search. Search crawlers stay on by default, and Google's training crawler (Google-Extended) is separate from its search crawler (Googlebot). Blocking training has no effect on ranking, and Google, Cloudflare and the Accountable companies have all said so in writing.
What can be blocked by default is training, and on pages with ads, agents. The agent case is the one to think about. When a customer asks ChatGPT "does this clinic take my insurance?", an agent fetches your page to answer. If that fetch is blocked, the answer comes from a competitor, a directory, or nowhere.
The four decisions
Whether or not you use Cloudflare, the change makes one thing clear: every website is now expected to state its AI policy. That comes down to four questions.
- Can search engines index us? Almost always yes. This includes ChatGPT search and Perplexity, which send people to your site.
- Can AI companies train on our content? Your call. Saying no does not affect search. Businesses with original content, pricing or expertise often choose no. Businesses that want to be known to AI assistants sometimes choose yes, because models learn about companies from training data.
- Can AI agents visit on a person's behalf? Usually yes. Blocking them means assistants can't answer questions about you.
- What must people and agents be told? If you use AI on your own site, such as a chatbot, several states already require a notice. Our state-by-state guide shows which ones.
The difficulty is that each answer lives in a different place: robots.txt rules, a Content-Signal line, Cloudflare's dashboard, a licence file, and a page humans can read. They drift apart, and most owners don't know what their site currently says.
Find out what your site says today
The free AI access check reads your site's robots.txt, home page headers and llms.txt, and reports what search engines, AI agents and AI training crawlers are allowed to do, crawler by crawler. It also shows whether you've published any of the newer signals. No sign-up.
Two things it will tell you that surprise most people:
- "Not stated" is the most common result. Nothing on the site asks AI companies to stay away, so well-behaved crawlers treat that as permission.
- Accidental blocking is real. A web designer pastes a "block all AI bots" file from a blog post and the site quietly drops out of ChatGPT search and Perplexity, with no change to its Google ranking to warn anyone.
The longer explanation of every signal a site can publish is in our guide to the September 15 change.
Frequently asked questions
Does this affect my site if I don't use Cloudflare?
Not directly. The default changes only happen on Cloudflare. But the three categories, the Content Signals line and the robots.txt rules work anywhere, and the AI companies that split their crawlers did so for the whole web, not just for Cloudflare customers.
Will blocking AI training hurt my Google ranking?
No. Google's training crawler is separate from its search crawler, and Google has said blocking it does not affect search. Cloudflare's Disallow AI Training setting is built on the same separation.
Is a robots.txt rule actually enforced?
It's a request. The large AI companies honour it, and that covers most training traffic. Crawlers that don't identify themselves ignore it. Real blocking happens at your host or at Cloudflare, not in the file, which is why the file and the host setting should say the same thing.
What should a small business do this week?
Run the free check, decide the four questions, and make sure your site actually says what you decided. If you're on Cloudflare, the AI Crawl Control section of the dashboard applies the answers for you.
Not legal advice. Cloudflare's settings and the standards mentioned here change; the linked sources are the authority.
