Cloudflare's Content Signals Policy defines the Content-Signal directive for robots.txt, giving site operators a machine-readable way to state downstream use preferences after a web crawler accesses a page. The three primary signals regulate distinct downstream operations: search governs traditional search indexing, ai-input governs retrieval-augmented generation and real-time grounding, and ai-train governs model training and fine-tuning.
Key facts
- Cloudflare introduced the Content Signals Policy on 2025-09-24 to address crawler data collection and AI scraping.
- Cloudflare engineers Michael Tremante and Leah Romm authored Internet-Draft
draft-romm-aipref-contentsignals-00; the IETF Datatracker records version 00 on 2025-10-01, and the draft expired on 2026-04-04. - The policy defines three signals,
search,ai-input, andai-train, evaluated via three-state logic:=yes(allowed),=no(disallowed), or omitted (no preference asserted). - While standard
AllowandDisallowdirectives control network retrieval access under RFC 9309,Content-Signalgoverns downstream data usage after access is granted. - Cloudflare's managed
robots.txtdefaults toContent-signal: search=yes, ai-train=no, use=referencewhile omittingai-inputto avoid assuming publisher preference. - Restrictions stated in the policy comments declare an explicit reservation of rights under Article 4 of European Union Directive 2019/790.
How Content Signals work
The Content-Signal directive appears inside a robots.txt record block alongside a target User-agent. It specifies comma-delimited key-value pairs that govern how automated crawlers may handle collected assets:
- `search`: Authorizes building a search index and returning search results, such as hyperlinks and short extracts. The specification explicitly states that
searchdoes not permit generating AI-synthesized search summaries. - `ai-input`: Authorizes ingesting content into AI models for real-time inference, including retrieval-augmented generation (RAG), conversational grounding, and real-time generative search answers.
- `ai-train`: Authorizes using content for training or fine-tuning machine learning models.
The specification uses three-state logic. Assigning =yes explicitly grants permission for that use. Assigning =no explicitly forbids that use. If a parameter is omitted, the publisher neither grants nor restricts permission for that category through the directive.
Cloudflare also tests an optional fourth parameter, use (or content-use), which sets retention limits:
use=immediate: The crawler may interact with the content in real time but cannot retain or store data.use=reference: The crawler may index, excerpt, and link back to the source.use=full: The crawler may summarize and reproduce the source material.
How Content Signals differ from Allow and Disallow
The Robots Exclusion Protocol (RFC 9309) uses Allow and Disallow to regulate crawler access to URL paths at the network transport layer. A crawler reading Disallow: /private/ is instructed not to make an HTTP request for that path. However, RFC 9309 provides no syntax to distinguish between an engine indexing a page for link referrals and an engine scraping text to train an LLM.
Content-Signal separates access from post-access exploitation:
Directive Type | Mechanism | Focus | Standard |
|---|---|---|---|
| Binary path match | Network fetch and crawl access | RFC 9309 |
| Key-value capability flags | Post-access data utilization | IETF |
Under this model, a site can serve Allow: / to permit public crawling while simultaneously setting ai-train=no to reserve intellectual property rights against model training. The accompanying policy comment frames this restriction under Article 4 of EU Directive 2019/790, which requires rightsholders to express machine-readable reservations against text and data mining exceptions.
Setting Content Signals on a site
Site operators can generate directives using the generator at contentsignals.org or add the line directly to robots.txt.
For example, https://insidetheloop.dev/robots.txt explicitly enables all three signals:
curl -s https://insidetheloop.dev/robots.txtUser-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=yes
Allow: /
# Disallow admin and API routes
Disallow: /_emdash/
Sitemap: https://insidetheloop.dev/sitemap.xmlIn this configuration:
search=yespermits traditional web crawlers to index pages for search queries.ai-input=yespermits AI agents to retrieve content at runtime for grounded summaries and RAG lookups.ai-train=yespermits AI labs to include published posts in future training corpora.
Cloudflare managed robots.txt and defaults
Cloudflare offers an automated managed robots.txt toggle in its dashboard under Security Settings. When enabled, Cloudflare prepends managed directives to any existing origin robots.txt or generates a standalone file.
The managed configuration injects the following block:
# BEGIN Cloudflare Managed content
User-Agent: *
Content-signal: search=yes, ai-train=no, use=reference
Allow: /
User-agent: Amazonbot
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: GPTBot
Disallow: /
User-agent: meta-externalagent
Disallow: /
# END Cloudflare Managed ContentCloudflare configures these specific defaults:
- `search=yes` and `ai-train=no`: Retains traditional search visibility while blocking AI training datasets.
- Omission of `ai-input`: Cloudflare intentionally leaves
ai-inputblank in managed deployments because the platform does not assume whether an operator wants to participate in real-time generative search answers. - Crawler-specific `Disallow` rules: Disallows eight known AI scrapers at the path level (
Amazonbot,Applebot-Extended,Bytespider,CCBot,ClaudeBot,Google-Extended,GPTBot, andmeta-externalagent). - Free plan fallback: Free plan domains without an origin
robots.txtfile and without managedrobots.txtenabled serve the explanatory Content Signals Policy comment block by default, without injecting machine-readable signal lines or access restrictions.
Cloudflare also documents a separate HTTP-header path for its Markdown for Agents product. In a 2026-07-13 changelog, Cloudflare said the product preserves an origin Content-Signal header and, when the origin sends none, adds Content-Signal: ai-train=yes, search=yes, ai-input=yes. That header behavior is separate from the robots.txt line described here.
Sources
- Giving users choice with Cloudflare's new Content Signals Policy: https://blog.cloudflare.com/content-signals-policy/ (read 2026-10-06)
- Cloudflare Gives Creators New Tool to Control Use of Their Content: https://www.cloudflare.com/press/press-releases/2025/cloudflare-gives-creators-new-tool-to-control-use-of-their-content/ (read 2026-10-06)
- robots.txt setting (Cloudflare Docs): https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/ (read 2026-10-06)
- Vocabulary For Expressing Content Signals (draft-romm-aipref-contentsignals-00): https://datatracker.ietf.org/doc/html/draft-romm-aipref-contentsignals-00 (read 2026-10-06)
- IETF Datatracker draft-romm-aipref-contentsignals: https://datatracker.ietf.org/doc/draft-romm-aipref-contentsignals/ (read 2026-10-06)
- Content Signals Generator: https://contentsignals.org/ (read 2026-10-06)
- Inside the Loop robots.txt: https://insidetheloop.dev/robots.txt (read 2026-10-06)
- Origin Content Signals for Markdown for Agents: https://developers.cloudflare.com/changelog/post/2026-07-13-markdown-for-agents-header-preservation/ (read 2026-10-06)
Last verified: 2026-10-06.