Open the robots.txt of a site that runs through Cloudflare and there is a fair chance you will find a line that Google’s own robots.txt documentation never mentions. It looks like this.
User-Agent: *
Content-signal: search=yes, ai-train=no, use=reference
Allow: /
That is a Content Signal. Content Signals is Cloudflare’s vocabulary for telling a crawler what it may do with your pages after it fetches them. It names three uses, search, AI answers and model training, and lets you say yes or no to each.
Cloudflare launched it on September 24, 2025, with its managed robots.txt already switched on for more than 3.8 million domains. The question people ask about it is simple. Does it do anything? Mostly, no. But Cloudflare’s managed robots.txt does not ship the line alone, and one of the lines it adds underneath does change where your site can appear in AI answers.
- It is a preference, not a block.
Content-Signalsits inside a robots.txt group and says yes or no tosearch,ai-input(feeding AI answers) andai-train(training models). Leave a value out and you neither grant nor refuse it. - Google does not read it. Google’s robots.txt spec supports four fields and ignores the rest. OpenAI, Anthropic and Perplexity do not mention it in their crawler documentation.
- Cloudflare’s managed file does more than one line. Under the signal it disallows eight named crawlers, including GPTBot, ClaudeBot and Google-Extended, in the syntax the big AI crawlers document reading.
- One of those lines costs you AI visibility. Disallowing Google-Extended takes your pages out of grounding in Gemini Apps. It does not touch Google Search or AI Overviews, and disallowing GPTBot does not touch ChatGPT search.
- Measure before you opt out of AI answers. Enforcing a no, by disallowing the crawlers that feed ChatGPT, Gemini, Claude and Perplexity, costs you the citations you hold there today. ContextBolt SEO, the SEO toolkit that lives inside the agent you already use, checks all four live. 7-day free trial, then $35 a month.
Where you are seeing this
If you found the line rather than searched for it, it probably came from one of these places.
- Your own robots.txt: after someone turned on the Cloudflare setting called Set your preference to block training in robots.txt, under Security Settings. Cloudflare prepends its block to the file you already had, or creates one if you had none.
- A Free plan domain with no robots.txt of its own: Cloudflare serves the policy text there, a block of
#comments defining the three signals, with no yes or no values at all. - Search Console’s robots.txt report: Cloudflare’s documentation warns it may flag the line as
Syntax not understood, and says Cloudflare has seen no effect on crawl rates or SEO from those reports. - The generator at contentsignals.org: it writes the block for you from four preset policies, from Disallow All to allowing everything.
- A Stack Overflow or Reddit thread: usually titled some version of “what does
search=yes,ai-train=noactually do”. You are in the right place.
What the three signals mean
Every definition below comes from the policy text Cloudflare publishes at contentsignals.org, shortened, followed by what each one means for a site owner.
search: building a search index and returning links and short excerpts. Cloudflare’s text explicitly excludes AI-generated search summaries. For you, this is the classic blue link.ai-input: feeding content into AI models at answer time, such as retrieval-augmented generation or grounding. For you, this is being read to answer somebody’s question right now.ai-train: training or fine-tuning AI models. For you, this is becoming part of a model.
Each value is yes or no, and the policy spells out a third state. A yes means a crawler may collect your content for that use. A no means it may not. A signal you leave out means you “neither grant nor restrict” that use through Content Signals, which is not the same as saying yes.
The most interesting line in that list is the exclusion under search. In Cloudflare’s vocabulary an AI summary on a results page is ai-input, not search. That is exactly the split Google does not offer. AI Overviews are fed by Googlebot, the same crawler as your blue link, and Google’s controls for AI features are nosnippet, data-nosnippet, max-snippet and noindex, every one of which also shrinks or removes your normal listing. The vocabulary describes a preference that Google offers no control for.
How the line is written
The signal lives inside a User-Agent group, next to Allow and Disallow. Values are comma-separated pairs. Give one crawler its own group to set a different preference for it, and put a path first to scope the signal to part of your site. This is the path form from Cloudflare’s own examples.
User-Agent: *
Content-Signal: /blog/ ai-train=no, search=yes, ai-input=no
Allow: /blog/
You will see it spelled Content-Signal and Content-signal. Google’s spec says robots.txt field names are case-insensitive, and the two spellings are meant as the same line.
One trap is worth knowing before you write your own. Google’s robots.txt spec says only one group is valid for a given crawler, the most specific one that matches it, and the rest are ignored. So a crawler with its own group never reads the signal in your User-Agent: * group. If you write a signal for everyone and then give one bot its own rules, say Googlebot with a Disallow: /private/, repeat the signal inside that bot’s group.
The fourth field, use=
On July 1, 2026 Cloudflare started testing a fourth field, use, which describes what a crawler may keep after it visits. It takes three values, from least to most permissive.
use=immediate: interact, but store and reuse nothing.use=reference: index, excerpt and link back.use=full: summarize and reproduce.
Every site that already had the managed robots.txt switched on got use=reference added to its line. That is why the example at the top of this page ends the way it does. No AI company documents reading it.
What Cloudflare writes into your file
Now the rest of the file. This is the managed block from Cloudflare’s robots.txt documentation, last updated August 3, 2026, with the comment lines removed.
User-Agent: *
Content-signal: search=yes, ai-train=no, use=reference
Allow: /
User-agent: Amazonbot
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: GPTBot
Disallow: /
User-agent: meta-externalagent
Disallow: /
Count the groups. One carries a Content Signal. Eight are plain Disallow rules for named crawlers, written in the syntax that OpenAI, Google and Anthropic all document honoring. The signal is the headline. The eight groups are the working part.
Three more details in that file are worth reading closely.
ai-input is not set: Cloudflare’s default says yes to search and no to training, and says nothing about AI answers. Under the policy’s own rule, that neither grants nor refuses them.
The comment block claims legal weight: the # lines above the rules state that any restriction you express is an express reservation of rights under Article 4 of the EU’s 2019 copyright directive, the provision that lets rights holders opt out of commercial text and data mining. Read that way, the policy is a rights notice written in robots.txt. The other machine-readable reservations, including the W3C’s TDMRep file, are covered in ai.txt, and this post does not re-argue them.
The spec site hedges that claim itself: contentsignals.org says in plain words that courts and regulators may conclude robots.txt files do not impose enforceable legal obligations. The policy is released under a CC0 license, so anyone can copy it into their own file.
Who honors Content Signals today
This is the question that decides whether the line is worth your time, so it was checked against each company’s own documentation on September 23, 2026, not against anyone’s blog.
Google: its robots.txt specification lists the fields it supports, user-agent, allow, disallow and sitemap, and says other fields are not supported. In July 2026 Google’s John Mueller answered a thread on r/TechSEO about the directive, and Search Engine Roundtable reported his reply. As far as he knew, he wrote, it has “no effects whatsoever for any crawler or LLM”. He added that it was made up by a CDN and adds maintenance to your file.
OpenAI, Anthropic and Perplexity: none of their crawler pages mention Content Signals. OpenAI’s bots page, Anthropic’s crawler page and Perplexity’s crawler page all describe control the same way, through named user-agent tokens in robots.txt.
Cloudflare itself: this is where the signal has teeth, and only on Cloudflare’s network. In the same July 2026 announcement Cloudflare said it is starting to track content use for every bot in its BotBase directory, and that a bot found abusing these signals loses its Verified status and is no longer allowed. That is enforcement by the CDN in front of your site, not by the AI company reading your file.
The standard that may replace it: the IETF’s AI Preferences working group is drafting the standards-track version of the same idea, and it uses different words. Its vocabulary draft, last revised September 13, 2026, names three categories, train-ai, ai-use and search, set to y or n. Its companion draft attaches them to robots.txt as a Content-Usage rule, and its own example of a training opt-out reads Content-Usage: train-ai=n. Both are still Internet-Drafts. The IETF search category also allows training a model that is used only for search, which Cloudflare’s definitions do not address.
So if you write a Content Signal today, write it knowing two things. The likely future is a renamed version of the same line. And what enforces your preference on the day is the user-agent rule beside it.
Does Content-Signal affect SEO?
This is the question a Reddit thread on r/CloudFlare asks outright, and it has two answers, one for the signal and one for the file it arrives in.
The signal line does not. Google ignores fields it does not support, so Content-signal changes nothing about how Googlebot crawls or ranks you. Cloudflare’s documentation says the same from the other side, that it has seen no impact on crawl rates or SEO when Search Console flags the line.
The Disallow lines do not touch Google Search either. Google’s crawler documentation says Google-Extended does not affect a site’s inclusion in Google Search and is not a ranking signal. But several of those lines do change what happens in AI answers outside Google Search.
| Line in the managed file | Google Search and AI Overviews | AI answers everywhere else |
|---|---|---|
| The signal line | Nothing. Google does not read the field | Nothing any vendor documents |
Google-Extended disallowed | Nothing, per Google | Your pages stop being used for Gemini training and for grounding answers in Gemini Apps and Vertex AI |
GPTBot disallowed | Nothing | Out of OpenAI training. ChatGPT search runs on OAI-SearchBot, which the file does not name |
ClaudeBot disallowed | Nothing | Out of Anthropic training. Claude-SearchBot and Claude-User are not in the file |
The other five groups, Amazonbot, Applebot-Extended, Bytespider, CCBot and meta-externalagent, are covered one by one in our AI crawlers list, with what is documented about each one.
Here is my take, and it is not a comfortable one for anyone who switched the setting on to feel protected. The line the whole internet argues about has no documented effect. A line that can cost you an AI citation is User-agent: Google-Extended, and it arrives in the same click as the signal, whether you read it or not. If Gemini naming your product matters to your business, that one group deserves more thought than the signal above it.
ai-input is the value that would cost you
ai-train=no is the easy call for most sites. Training takes your pages and gives nothing back in the moment, and there is no link in a model’s weights.
ai-input is different. It covers the use that produces a citation. An AI answer that grounds itself on your page is the answer that can name you and link to you. Write ai-input=no and you are asking to be left out of exactly the answers that send people your way.
Today no vendor reads the value, so writing it costs nothing and does nothing beyond the rights notice. The versions a crawler does obey are per vendor, and they look like this.
- Stay in classic search,
search=yes: allow Googlebot and Bingbot, which is the default. - Stay out of model training,
ai-train=no: disallow GPTBot, ClaudeBot and CCBot. Google-Extended too, but read the next paragraph first. - Stay out of AI answers,
ai-input=no: disallow OAI-SearchBot, Claude-SearchBot, PerplexityBot and Google-Extended. For AI Overviews there is onlynosnippetornoindex, and both also cut your normal listing. - Stay in AI answers,
ai-input=yes: allow those search crawlers and leave Google-Extended out of your file.
Google has no token for training alone. Google-Extended covers Gemini training and Gemini grounding together, so on Google’s side ai-train=no with ai-input=yes cannot be enforced through robots.txt. You pick one outcome for both. OpenAI, Anthropic and Perplexity do separate the jobs. OpenAI’s page says a site can disallow GPTBot and still appear in ChatGPT search through OAI-SearchBot, and Perplexity says PerplexityBot is not used to crawl content for AI foundation models.
One more gap. The fetchers that act when a user asks about a specific page, ChatGPT-User and Perplexity-User, sit partly outside robots.txt. OpenAI says its robots.txt rules might not cover visits a user starts, and Perplexity says its user fetcher generally ignores them. Whether any of this is worth blocking in the first place is a separate decision, and we take it apart in should you block AI crawlers.
What to put in your file
The position most publishers actually hold is yes to search, yes to AI answers, no to training. Written so that the signal states it and the user-agent rules enforce it, the file looks like this.
Copy this block
User-Agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Allow: /
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
Disallow: /
A few notes before you paste it.
The signal lives in the * group, and that is enough here. GPTBot, ClaudeBot and CCBot read their own group and skip the * one, but they are told to stay out entirely, so there is nothing for a signal to add. Give any crawler a group with partial rules and you would repeat the signal there, per the one-group rule.
Stacking user-agent lines is standard. Google’s spec shows several user-agent lines sharing one set of rules. It keeps the file short and the intent easy to read.
Google-Extended is left out deliberately. Add a User-agent: Google-Extended group with Disallow: / only if keeping out of Gemini training matters more to you than being cited by Gemini, because the token refuses both.
If Cloudflare manages your robots.txt, you already have most of this. Open https://example.com/robots.txt on your own domain and look for User-agent: Google-Extended. If it is there and you care about Gemini citations, that is a decision worth making on purpose rather than by default.
The difference between this file and llms.txt, the map that tells AI readers what to read rather than what they may take, is covered in llms.txt vs robots.txt. The older training opt-out that shares the space, ai.txt, has its own post, linked above.
Check who cites you before you say no
Every rule on this page that keeps a crawler out of AI answers has a price, and it is paid in citations. Enforcing a no to AI answers, by disallowing the crawlers that feed them, costs you whatever ChatGPT, Gemini, Claude and Perplexity say about you today. Nobody can tell you that number from a robots.txt file. You have to ask the engines.
Copy this prompt
Read example.com and write the three questions a buyer would ask
an AI before choosing a product like ours. Ask ChatGPT, Gemini,
Claude and Perplexity each question. For every answer, tell me
whether example.com is named, cited as a source, or absent, and
list every source the engine cited.
Twelve live checks at two credits each is 24 credits, out of the 1,000 a month the plan includes. Read the result by engine.
Gemini cites you: think hard before the Google-Extended line. That is the citation it removes.
ChatGPT or Claude cite you and you want out of training: you can have both. Disallow GPTBot and ClaudeBot, allow OAI-SearchBot and Claude-SearchBot.
Nobody cites you: the signal choice is academic for now. The work is on the page, and how AI assistants choose citations is where to start. If an engine reads your page and names a rival anyway, that has its own fix, in cited but not recommended.
ContextBolt SEO is ours, so weigh the bias. This is the job it was built for. Ask in plain English and it asks each engine live, then reports whether you were named, cited or absent, with every source the engine used. Ask how often your brand appears across ChatGPT and Google’s AI answers and it pulls that too, estimated from a large index of answers, decision-useful and directionally right. AI engines do not answer the same question the same way every time, so a live check tells you whether you showed up that time. Every result saves automatically to your SEO Dashboard, dated, so you can check again a month after you change your robots.txt and see what moved.
So write the Content Signal if it states your position, because the line itself costs nothing and it is a rights notice. Put the rules crawlers actually read beside it, and repeat the signal in any group that is not a full Disallow. Then measure the thing no line in robots.txt can report, which is whether the answer engines cite you. That measurement is one question away on the ContextBolt SEO trial.