Which sites search engines surface for 50 AI-agent questions: our October 2026 measurement

For 50 AI-agent questions, Google gave community sites 30.4% of candidate rows; Perplexity gave them 1.6%, with more vendor documentation and GitHub.

Update (2026-10-07): Recomputed every measurement from the original 500-row CSV and separated all-domain totals from the independent-publisher subset.

On 2026-10-07, Google Serper and the Perplexity Search API returned different candidate pools for the same 50 AI-agent questions. Google gave community sites 30.4% of its 250 rows, including Reddit at 17.6%; Perplexity gave them 1.6%. Perplexity returned vendor and official sites at 30.0% and GitHub rows at 8.8%.

Key facts

  • On 2026-10-07, 50 questions produced 500 candidate rows: 250 from Google Serper and 250 from the Perplexity Search API.
  • The 500 rows contained 235 distinct domains. Google used 124, Perplexity used 148, and 37 domains appeared in both engines, or 15.7% of all distinct domains.
  • The 500 rows contained 472 distinct URLs. The engines shared 27 exact URLs, or 5.7% of the distinct-URL pool.
  • Reddit captured 44 Google rows, or 17.6%, and 0 Perplexity rows.
  • The Vendor Docs & Official Sites category represented 18.8% of Google rows and 30.0% of Perplexity rows.
  • Independent Blogs & Tech Publishers captured 236 rows, or 47.2%, across 188 distinct domains. All categories together contained 235 distinct domains.
  • insidetheloop.dev appeared 0 times across the 500 candidate rows.

What does this October 2026 retrieval benchmark measure?

This measurement captures candidate retrieval pools returned by search layers that answer engines query. It does not measure the final citations or prose produced by a chat model.

Benchmark parameter

Specification

Run date

2026-10-07

Query count

50 practitioner questions across 6 archetypes

Google search layer

Top 5 organic results per query via Serper (250 rows)

Perplexity search layer

Top 5 results per query via the Perplexity Search API (250 rows)

Dataset size

500 ranked rows across 235 distinct domains

Measurement scope

Upstream candidate retrieval pools, not LLM answer synthesis

The category labels use the fixed domain sets in the analysis script. The independent category is the remainder after named community platforms, GitHub, vendor and official sites, social and forum sites, and academic domains are removed.

A freshness search on 2026-10-07 found a separate 2026 AI-search study with a different corpus. It does not change these fixed measurements. AI Search Statistics for 2026 is included as context in the Sources section.

Which domains dominate Google organic results versus Perplexity Search API?

Google organic search and the Perplexity Search API produced different distributions for the same developer prompts.

Source category

Google Serper (N=250)

Perplexity API (N=250)

Total pool (N=500)

Independent blogs & publishers

100 (40.0%)

136 (54.4%)

236 (47.2%)

Vendor docs & official portals

47 (18.8%)

75 (30.0%)

122 (24.4%)

Core community (Reddit/YouTube/Medium/dev.to)

76 (30.4%)

4 (1.6%)

80 (16.0%)

GitHub (repos & docs)

13 (5.2%)

22 (8.8%)

35 (7.0%)

Social & forums

13 (5.2%)

5 (2.0%)

18 (3.6%)

Academic & preprints

1 (0.4%)

8 (3.2%)

9 (1.8%)

The community-platform counts were:

Platform

Google organic results

Perplexity Search API results

Combined count

reddit.com

44 (17.6%)

0 (0.0%)

44 (8.8%)

youtube.com

18 (7.2%)

0 (0.0%)

18 (3.6%)

medium.com

9 (3.6%)

0 (0.0%)

9 (1.8%)

dev.to

5 (2.0%)

4 (1.6%)

9 (1.8%)

Reddit was Google's most frequent domain, with 44 of 250 rows. Perplexity returned no Reddit, YouTube, or Medium rows. github.com was Perplexity's most frequent single domain, with 20 rows, or 8.0%. The broader GitHub category has 22 rows, or 8.8%, because it also includes docs.github.com. arxiv.org appeared 7 times in Perplexity rows, or 2.8%.

How does retrieval shift across different question kinds?

Retrieval mixes changed by query intent across the six question archetypes in the dataset:

Question archetype

Google community share

Perplexity community share

Perplexity vendor + GitHub share

best-tool (N=90)

40.0% (18/45)

2.2% (1/45)

17.8% (8/45)

comparisons (N=90)

40.0% (18/45)

2.2% (1/45)

0.0% (0/45)

what-changed (N=80)

30.0% (12/40)

0.0% (0/40)

80.0% (32/40)

how-it-works (N=80)

27.5% (11/40)

2.5% (1/40)

55.0% (22/40)

troubleshooting (N=80)

25.0% (10/40)

2.5% (1/40)

45.0% (18/40)

recent-news (N=80)

17.5% (7/40)

0.0% (0/40)

42.5% (17/40)

For comparisons, 43 of 45 Perplexity rows, or 95.6%, were in the independent category. One row came from LinkedIn and one from dev.to, so the broader remainder after community, vendor, and GitHub rows was 44 of 45, or 97.8%. For what-changed queries, Perplexity assigned 25 of 40 rows to vendor sources, or 62.5%, and 7 to GitHub, or 17.5%. Together, those categories held 80.0% of the rows. Google gave 30.0% of its version-query rows to community platforms.

Where did search candidate pools return the weakest results?

Three questions exposed clear topic mismatches in the candidate pools:

  1. Aider concurrent Git conflicts (Q30): Perplexity's five rows were general Git conflict documentation or GitHub pages, not Aider-specific pages. Google included Aider issue #800 alongside Stack Overflow, Reddit, and a GitHub community discussion.
  2. Aider versus Windsurf multi-file refactoring (Q16): No result title described a direct empirical Aider-versus-Windsurf diff benchmark. Perplexity's first result was joinleland.com; Google mixed Aider, Cursor, and Windsurf comparisons with Reddit and dev.to pages.
  3. Gemini 2.0 Flash tool latency (Q25): Perplexity included two iTechGuides pages, including one whose title referred to a 2026 shutdown. Google returned a Gemini API changelog and newer Gemini pages. The pool did not produce a focused latency measurement.

What does the zero for insidetheloop.dev show?

It shows that none of the 500 candidate rows used insidetheloop.dev. The CSV cannot establish whether that absence came from crawling, ranking, query wording, or another part of retrieval.

The related posts llms.txt and llms-full.txt: what they are, who reads them, and how to generate them and Cloudflare Content Signals in robots.txt: what search, ai-input, and ai-train mean document the site's crawler-facing controls. This benchmark does not test whether those controls change rankings. How AI agents fetch web pages: user agents, IP ranges, and fetch origins covers the separate fetch step.

The retrieved independent pages also show useful patterns to test, but they do not prove that any one pattern caused a ranking:

Sources

Last verified: 2026-10-07.

Spotted an outdated or wrong claim? Agents can report it with evidence throughPOST /api/feedback; an editor checks every report. See llms.txt for the agent API.