Hosted AI bots fetch web pages from cloud datacenter IP ranges published in machine-readable JSON feeds or verified via reverse DNS. Terminal coding agents like Claude Code WebFetch originate requests directly from the user's workstation, using Claude-User while bypassing vendor IP allowlists. A curl test with a spoofed user agent fails Cloudflare bot checks because Cloudflare validates source IP ranges, reverse DNS, and RFC 9421 cryptographic signatures rather than trusting user-agent headers alone.
Key facts
- Hosted crawlers (OpenAI GPTBot, OAI-SearchBot, ChatGPT-User; Anthropic ClaudeBot, Claude-SearchBot, Claude-User; PerplexityBot, Perplexity-User) publish CIDR prefixes in JSON feeds on their primary domains.
- Search bots Googlebot and Microsoft bingbot originate from cloud IP lists and validate through two-way reverse and forward DNS lookups ending in
googlebot.comandsearch.msn.com. - Google-Extended is a robots.txt access control token, not an HTTP user-agent header; requests to sites blocking Google-Extended arrive under standard Googlebot user-agent strings.
- Claude Code WebFetch executes HTTP requests locally on developer machines with
Accept: text/markdown, text/html, */*and user agentClaude-User (claude-code/<version>; ...), sending only target hostnames toapi.anthropic.comfor safety filtering. - Gemini CLI
web_fetchretrieves content server-side via the Gemini APIurlContextservice, falling back to direct local fetches if API retrieval fails. - OpenAI Codex CLI web search operates server-side in hosted environments, defaulting to cached index retrieval to prevent prompt injection.
- Cloudflare verifies bot authenticity using RFC 9421 HTTP Message Signatures (Web Bot Auth), published IP allowlists, and reverse DNS; testing with
curl -Atriggers impersonation blocks against verified bot rules.
Crawler and agent fetch reference
Crawler or Agent | Role | Origin | User-Agent Header | Verification Method |
|---|---|---|---|---|
GPTBot | Model training | OpenAI servers |
|
|
OAI-SearchBot | Search index | OpenAI servers |
|
|
ChatGPT-User | User actions | OpenAI servers |
|
|
ClaudeBot | Model training | Anthropic servers |
|
|
Claude-SearchBot | Search index | Anthropic servers |
|
|
Claude-User | User actions | Anthropic servers |
|
|
PerplexityBot | Search index | Perplexity servers |
|
|
Perplexity-User | User actions | Perplexity servers |
|
|
Googlebot | Search engine | Google servers |
| rDNS |
Google-Extended | robots.txt token | Google servers | None (arrives under | robots.txt token only |
Google-Agent | User-initiated AI actions (e.g., Project Mariner) | Google servers |
| rDNS |
bingbot | Search engine | Microsoft servers |
| rDNS |
Claude Code WebFetch | CLI coding agent | Local machine |
| Local IP; preflight to |
Gemini CLI web_fetch | CLI coding agent | Google API / Local | Gemini API / local client HTTP request | Google API; local IP on fallback |
OpenAI Codex CLI | CLI coding agent | OpenAI hosted | Hosted web retrieval agent | Hosted tool; cached index by default |
Who fetches from where: vendor datacenters versus developer workstations
Web traffic generated by AI systems follows two network architectures: centralized vendor infrastructure and distributed client machines.
Centralized crawlers run from vendor datacenters. OpenAI publishes IP ranges for GPTBot, OAI-SearchBot, and ChatGPT-User. Anthropic publishes 28 CIDR prefixes in https://claude.com/crawling/bots.json (as of 2026-10-02). Perplexity hosts IP lists for PerplexityBot and Perplexity-User. Google and Microsoft rely on cloud ranges with forward-confirmed reverse DNS (FCrDNS). On March 20, 2026, Google added Google-Agent — a user-triggered fetcher for AI agents such as Project Mariner — with its own IP range file (user-triggered-agents.json). Firewalls identify these services by matching IP endpoints.
In contrast, terminal coding agents execute fetches from the developer's local machine:
- Claude Code WebFetch: Before opening a connection, Claude Code sends the target hostname to
https://api.anthropic.comfor safety blocklist verification. Once cleared, the developer machine issues the HTTP request directly from its local IP address. The request sendsAccept: text/markdown, text/html, */*and user-agentClaude-User (claude-code/2.1.258; +https://support.anthropic.com/)(builds prior to version 2.1.258 sentaxios/1.8.4). Because traffic originates from client IP addresses, it bypasses WAF rules allowlisting Anthropic datacenter ranges. - Gemini CLI web_fetch: Gemini CLI uses a hybrid design. It routes URLs through the server-side Gemini API
urlContextservice. If API retrieval fails, it falls back to fetching directly from the local workstation with connection pinning against internal subnets. - OpenAI Codex CLI: Codex executes web search as a hosted server tool. By default, Codex runs in
cachedmode against an OpenAI search index. Inlivemode (web_search = "live"), retrieval runs on hosted OpenAI servers rather than local client sockets.
Why a spoofed-UA curl test fails Cloudflare bot verification
A common testing shortcut is running curl with a spoofed User-Agent header:
curl -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" https://example.comThis test fails to reproduce Cloudflare's verified bot checks.
Cloudflare does not grant verified bot status based on the User-Agent header alone. Instead, Cloudflare enforces verified bot rules through three verification mechanisms:
- Web Bot Auth (RFC 9421): Agents cryptographically sign requests using HTTP Message Signatures. The client sends
Signature,Signature-Input, andSignature-Agentheaders containing public key thumbprints and timestamps, validated at the edge against author public keys. - IP range validation: For crawlers like GPTBot, ClaudeBot, and PerplexityBot, Cloudflare matches connecting IP addresses against verified CIDR feeds published by each vendor.
- Forward-confirmed reverse DNS: For search engines like Googlebot and bingbot, Cloudflare verifies that reverse DNS resolves to
googlebot.comorsearch.msn.comand matches the forward lookup.
When a developer runs curl -A from an arbitrary workstation or cloud VM, the request presents a crawler user-agent string but fails IP validation, rDNS verification, and cryptographic signature checks. Cloudflare flags it as an unverified bot or spoofed bot and applies Bot Fight Mode challenges (cf-mitigated: challenge) or 403 blocks.
To verify access, inspect Cloudflare Security Events and AI Crawl Control analytics, or test through native agent interfaces.
Sources
- Overview of OpenAI Crawlers: https://platform.openai.com/docs/bots (read 2026-10-06)
- Does Anthropic crawl data from the web, and how can site owners block the crawler?: https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler (read 2026-10-06)
- Anthropic Crawling IP List: https://claude.com/crawling/bots.json (read 2026-10-06)
- Claude Code Tools Reference: https://code.claude.com/docs/en/tools-reference (read 2026-10-06)
- Claude Code Data Usage: https://code.claude.com/docs/en/data-usage (read 2026-10-06)
- How Claude Code Web Fetch Works: Your IP, Their Blocklist: https://www.picklog.cc/blog/claude-code-web-fetch (read 2026-10-06)
- Perplexity Crawlers Documentation: https://docs.perplexity.ai/docs/resources/perplexity-crawlers (read 2026-10-06)
- Google Common Crawlers and Fetchers: https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers (read 2026-10-06)
- Google User-Triggered Fetchers (Google-Agent): https://developers.google.com/crawling/docs/crawlers-fetchers/google-agent (read 2026-10-06)
- How to Verify Bingbot: https://webmasterid.com/search-bots/how-to-verify-bingbot (read 2026-10-06)
- Gemini CLI Web Fetch Tool: https://geminicli.com/docs/tools/web-fetch/ (read 2026-10-06)
- Codex CLI Web Search Configuration: https://codex.danielvaughan.com/2026/05/09/codex-cli-web-search-configuration-cached-live-domain-allow-lists-prompt-injection-defence/ (read 2026-10-06)
- Cloudflare Verified Bots: https://developers.cloudflare.com/bots/concepts/bot/verified-bots/ (read 2026-10-06)
- Forget IPs: using cryptography to verify bot and agent traffic: https://blog.cloudflare.com/web-bot-auth/ (read 2026-10-06)
Last verified: 2026-10-06.