How AI agents fetch web pages: user agents, IP ranges, and fetch origins

AI crawlers fetch from vendor datacenters using published IP ranges, while CLI coding agents fetch directly from developer workstations via local IPs.

Hosted AI bots fetch web pages from cloud datacenter IP ranges published in machine-readable JSON feeds or verified via reverse DNS. Terminal coding agents like Claude Code WebFetch originate requests directly from the user's workstation, using Claude-User while bypassing vendor IP allowlists. A curl test with a spoofed user agent fails Cloudflare bot checks because Cloudflare validates source IP ranges, reverse DNS, and RFC 9421 cryptographic signatures rather than trusting user-agent headers alone.

Key facts

  • Hosted crawlers (OpenAI GPTBot, OAI-SearchBot, ChatGPT-User; Anthropic ClaudeBot, Claude-SearchBot, Claude-User; PerplexityBot, Perplexity-User) publish CIDR prefixes in JSON feeds on their primary domains.
  • Search bots Googlebot and Microsoft bingbot originate from cloud IP lists and validate through two-way reverse and forward DNS lookups ending in googlebot.com and search.msn.com.
  • Google-Extended is a robots.txt access control token, not an HTTP user-agent header; requests to sites blocking Google-Extended arrive under standard Googlebot user-agent strings.
  • Claude Code WebFetch executes HTTP requests locally on developer machines with Accept: text/markdown, text/html, */* and user agent Claude-User (claude-code/<version>; ...), sending only target hostnames to api.anthropic.com for safety filtering.
  • Gemini CLI web_fetch retrieves content server-side via the Gemini API urlContext service, falling back to direct local fetches if API retrieval fails.
  • OpenAI Codex CLI web search operates server-side in hosted environments, defaulting to cached index retrieval to prevent prompt injection.
  • Cloudflare verifies bot authenticity using RFC 9421 HTTP Message Signatures (Web Bot Auth), published IP allowlists, and reverse DNS; testing with curl -A triggers impersonation blocks against verified bot rules.

Crawler and agent fetch reference

Crawler or Agent

Role

Origin

User-Agent Header

Verification Method

GPTBot

Model training

OpenAI servers

Mozilla/5.0 ... compatible; GPTBot/1.4; +https://openai.com/gptbot

openai.com/gptbot.json

OAI-SearchBot

Search index

OpenAI servers

Mozilla/5.0 ... Chrome/131.0.0.0 ... compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot

openai.com/searchbot.json

ChatGPT-User

User actions

OpenAI servers

Mozilla/5.0 ... compatible; ChatGPT-User/1.0; +https://openai.com/bot

openai.com/chatgpt-user.json

ClaudeBot

Model training

Anthropic servers

Mozilla/5.0 ... compatible; ClaudeBot/1.0; +claudebot@anthropic.com

claude.com/crawling/bots.json

Claude-SearchBot

Search index

Anthropic servers

Mozilla/5.0 ... compatible; Claude-SearchBot/1.0; +Claude-SearchBot@anthropic.com

claude.com/crawling/bots.json

Claude-User

User actions

Anthropic servers

Mozilla/5.0 ... compatible; Claude-User/1.0; +Claude-User@anthropic.com

claude.com/crawling/bots.json

PerplexityBot

Search index

Perplexity servers

Mozilla/5.0 ... compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot

perplexity.com/perplexitybot.json

Perplexity-User

User actions

Perplexity servers

Mozilla/5.0 ... compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user

perplexity.com/perplexity-user.json

Googlebot

Search engine

Google servers

Mozilla/5.0 ... Googlebot/2.1

rDNS *.googlebot.com + common-crawlers.json

Google-Extended

robots.txt token

Google servers

None (arrives under Googlebot header)

robots.txt token only

Google-Agent

User-initiated AI actions (e.g., Project Mariner)

Google servers

Mozilla/5.0 ... compatible; Google-Agent; ...

rDNS google.com + user-triggered-agents.json

bingbot

Search engine

Microsoft servers

Mozilla/5.0 ... bingbot/2.0

rDNS search.msn.com + bingbot.json

Claude Code WebFetch

CLI coding agent

Local machine

Claude-User (claude-code/<ver>; +https://support.anthropic.com/)

Local IP; preflight to api.anthropic.com

Gemini CLI web_fetch

CLI coding agent

Google API / Local

Gemini API / local client HTTP request

Google API; local IP on fallback

OpenAI Codex CLI

CLI coding agent

OpenAI hosted

Hosted web retrieval agent

Hosted tool; cached index by default

Who fetches from where: vendor datacenters versus developer workstations

Web traffic generated by AI systems follows two network architectures: centralized vendor infrastructure and distributed client machines.

Centralized crawlers run from vendor datacenters. OpenAI publishes IP ranges for GPTBot, OAI-SearchBot, and ChatGPT-User. Anthropic publishes 28 CIDR prefixes in https://claude.com/crawling/bots.json (as of 2026-10-02). Perplexity hosts IP lists for PerplexityBot and Perplexity-User. Google and Microsoft rely on cloud ranges with forward-confirmed reverse DNS (FCrDNS). On March 20, 2026, Google added Google-Agent — a user-triggered fetcher for AI agents such as Project Mariner — with its own IP range file (user-triggered-agents.json). Firewalls identify these services by matching IP endpoints.

In contrast, terminal coding agents execute fetches from the developer's local machine:

  1. Claude Code WebFetch: Before opening a connection, Claude Code sends the target hostname to https://api.anthropic.com for safety blocklist verification. Once cleared, the developer machine issues the HTTP request directly from its local IP address. The request sends Accept: text/markdown, text/html, */* and user-agent Claude-User (claude-code/2.1.258; +https://support.anthropic.com/) (builds prior to version 2.1.258 sent axios/1.8.4). Because traffic originates from client IP addresses, it bypasses WAF rules allowlisting Anthropic datacenter ranges.
  2. Gemini CLI web_fetch: Gemini CLI uses a hybrid design. It routes URLs through the server-side Gemini API urlContext service. If API retrieval fails, it falls back to fetching directly from the local workstation with connection pinning against internal subnets.
  3. OpenAI Codex CLI: Codex executes web search as a hosted server tool. By default, Codex runs in cached mode against an OpenAI search index. In live mode (web_search = "live"), retrieval runs on hosted OpenAI servers rather than local client sockets.

Why a spoofed-UA curl test fails Cloudflare bot verification

A common testing shortcut is running curl with a spoofed User-Agent header:

curl -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" https://example.com

This test fails to reproduce Cloudflare's verified bot checks.

Cloudflare does not grant verified bot status based on the User-Agent header alone. Instead, Cloudflare enforces verified bot rules through three verification mechanisms:

  1. Web Bot Auth (RFC 9421): Agents cryptographically sign requests using HTTP Message Signatures. The client sends Signature, Signature-Input, and Signature-Agent headers containing public key thumbprints and timestamps, validated at the edge against author public keys.
  2. IP range validation: For crawlers like GPTBot, ClaudeBot, and PerplexityBot, Cloudflare matches connecting IP addresses against verified CIDR feeds published by each vendor.
  3. Forward-confirmed reverse DNS: For search engines like Googlebot and bingbot, Cloudflare verifies that reverse DNS resolves to googlebot.com or search.msn.com and matches the forward lookup.

When a developer runs curl -A from an arbitrary workstation or cloud VM, the request presents a crawler user-agent string but fails IP validation, rDNS verification, and cryptographic signature checks. Cloudflare flags it as an unverified bot or spoofed bot and applies Bot Fight Mode challenges (cf-mitigated: challenge) or 403 blocks.

To verify access, inspect Cloudflare Security Events and AI Crawl Control analytics, or test through native agent interfaces.

Sources

Last verified: 2026-10-06.

Spotted an outdated or wrong claim? Agents can report it with evidence throughPOST /api/feedback; an editor checks every report. See llms.txt for the agent API.