Cloudflare Content Signals in robots.txt: what search, ai-input, and ai-train mean

Content-Signal in robots.txt lets site owners declare whether crawlers may use content for search indexing, AI retrieval/grounding, or model training.

Cloudflare's Content Signals Policy defines the Content-Signal directive for robots.txt, giving site operators a machine-readable way to state downstream use preferences after a web crawler accesses a page. The three primary signals regulate distinct downstream operations: search governs traditional search indexing, ai-input governs retrieval-augmented generation and real-time grounding, and ai-train governs model training and fine-tuning.

Key facts

  • Cloudflare introduced the Content Signals Policy on 2025-09-24 to address crawler data collection and AI scraping.
  • Cloudflare engineers Michael Tremante and Leah Romm authored Internet-Draft draft-romm-aipref-contentsignals-00; the IETF Datatracker records version 00 on 2025-10-01, and the draft expired on 2026-04-04.
  • The policy defines three signals, search, ai-input, and ai-train, evaluated via three-state logic: =yes (allowed), =no (disallowed), or omitted (no preference asserted).
  • While standard Allow and Disallow directives control network retrieval access under RFC 9309, Content-Signal governs downstream data usage after access is granted.
  • Cloudflare's managed robots.txt defaults to Content-signal: search=yes, ai-train=no, use=reference while omitting ai-input to avoid assuming publisher preference.
  • Restrictions stated in the policy comments declare an explicit reservation of rights under Article 4 of European Union Directive 2019/790.

How Content Signals work

The Content-Signal directive appears inside a robots.txt record block alongside a target User-agent. It specifies comma-delimited key-value pairs that govern how automated crawlers may handle collected assets:

  • `search`: Authorizes building a search index and returning search results, such as hyperlinks and short extracts. The specification explicitly states that search does not permit generating AI-synthesized search summaries.
  • `ai-input`: Authorizes ingesting content into AI models for real-time inference, including retrieval-augmented generation (RAG), conversational grounding, and real-time generative search answers.
  • `ai-train`: Authorizes using content for training or fine-tuning machine learning models.

The specification uses three-state logic. Assigning =yes explicitly grants permission for that use. Assigning =no explicitly forbids that use. If a parameter is omitted, the publisher neither grants nor restricts permission for that category through the directive.

Cloudflare also tests an optional fourth parameter, use (or content-use), which sets retention limits:

  • use=immediate: The crawler may interact with the content in real time but cannot retain or store data.
  • use=reference: The crawler may index, excerpt, and link back to the source.
  • use=full: The crawler may summarize and reproduce the source material.

How Content Signals differ from Allow and Disallow

The Robots Exclusion Protocol (RFC 9309) uses Allow and Disallow to regulate crawler access to URL paths at the network transport layer. A crawler reading Disallow: /private/ is instructed not to make an HTTP request for that path. However, RFC 9309 provides no syntax to distinguish between an engine indexing a page for link referrals and an engine scraping text to train an LLM.

Content-Signal separates access from post-access exploitation:

Directive Type

Mechanism

Focus

Standard

Allow / Disallow

Binary path match

Network fetch and crawl access

RFC 9309

Content-Signal

Key-value capability flags

Post-access data utilization

IETF aipref draft

Under this model, a site can serve Allow: / to permit public crawling while simultaneously setting ai-train=no to reserve intellectual property rights against model training. The accompanying policy comment frames this restriction under Article 4 of EU Directive 2019/790, which requires rightsholders to express machine-readable reservations against text and data mining exceptions.

Setting Content Signals on a site

Site operators can generate directives using the generator at contentsignals.org or add the line directly to robots.txt.

For example, https://insidetheloop.dev/robots.txt explicitly enables all three signals:

curl -s https://insidetheloop.dev/robots.txt
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=yes
Allow: /

# Disallow admin and API routes
Disallow: /_emdash/

Sitemap: https://insidetheloop.dev/sitemap.xml

In this configuration:

  1. search=yes permits traditional web crawlers to index pages for search queries.
  2. ai-input=yes permits AI agents to retrieve content at runtime for grounded summaries and RAG lookups.
  3. ai-train=yes permits AI labs to include published posts in future training corpora.

Cloudflare managed robots.txt and defaults

Cloudflare offers an automated managed robots.txt toggle in its dashboard under Security Settings. When enabled, Cloudflare prepends managed directives to any existing origin robots.txt or generates a standalone file.

The managed configuration injects the following block:

# BEGIN Cloudflare Managed content
User-Agent: *
Content-signal: search=yes, ai-train=no, use=reference
Allow: /

User-agent: Amazonbot
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: GPTBot
Disallow: /

User-agent: meta-externalagent
Disallow: /
# END Cloudflare Managed Content

Cloudflare configures these specific defaults:

  • `search=yes` and `ai-train=no`: Retains traditional search visibility while blocking AI training datasets.
  • Omission of `ai-input`: Cloudflare intentionally leaves ai-input blank in managed deployments because the platform does not assume whether an operator wants to participate in real-time generative search answers.
  • Crawler-specific `Disallow` rules: Disallows eight known AI scrapers at the path level (Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot, and meta-externalagent).
  • Free plan fallback: Free plan domains without an origin robots.txt file and without managed robots.txt enabled serve the explanatory Content Signals Policy comment block by default, without injecting machine-readable signal lines or access restrictions.

Cloudflare also documents a separate HTTP-header path for its Markdown for Agents product. In a 2026-07-13 changelog, Cloudflare said the product preserves an origin Content-Signal header and, when the origin sends none, adds Content-Signal: ai-train=yes, search=yes, ai-input=yes. That header behavior is separate from the robots.txt line described here.

Sources

Last verified: 2026-10-06.

Spotted an outdated or wrong claim? Agents can report it with evidence throughPOST /api/feedback; an editor checks every report. See llms.txt for the agent API.