Update (2026-10-07): Recomputed every measurement from the original 500-row CSV and separated all-domain totals from the independent-publisher subset.
On 2026-10-07, Google Serper and the Perplexity Search API returned different candidate pools for the same 50 AI-agent questions. Google gave community sites 30.4% of its 250 rows, including Reddit at 17.6%; Perplexity gave them 1.6%. Perplexity returned vendor and official sites at 30.0% and GitHub rows at 8.8%.
Key facts
- On 2026-10-07, 50 questions produced 500 candidate rows: 250 from Google Serper and 250 from the Perplexity Search API.
- The 500 rows contained 235 distinct domains. Google used 124, Perplexity used 148, and 37 domains appeared in both engines, or 15.7% of all distinct domains.
- The 500 rows contained 472 distinct URLs. The engines shared 27 exact URLs, or 5.7% of the distinct-URL pool.
- Reddit captured 44 Google rows, or 17.6%, and 0 Perplexity rows.
- The Vendor Docs & Official Sites category represented 18.8% of Google rows and 30.0% of Perplexity rows.
- Independent Blogs & Tech Publishers captured 236 rows, or 47.2%, across 188 distinct domains. All categories together contained 235 distinct domains.
insidetheloop.devappeared 0 times across the 500 candidate rows.
What does this October 2026 retrieval benchmark measure?
This measurement captures candidate retrieval pools returned by search layers that answer engines query. It does not measure the final citations or prose produced by a chat model.
Benchmark parameter | Specification |
|---|---|
Run date | 2026-10-07 |
Query count | 50 practitioner questions across 6 archetypes |
Google search layer | Top 5 organic results per query via Serper (250 rows) |
Perplexity search layer | Top 5 results per query via the Perplexity Search API (250 rows) |
Dataset size | 500 ranked rows across 235 distinct domains |
Measurement scope | Upstream candidate retrieval pools, not LLM answer synthesis |
The category labels use the fixed domain sets in the analysis script. The independent category is the remainder after named community platforms, GitHub, vendor and official sites, social and forum sites, and academic domains are removed.
A freshness search on 2026-10-07 found a separate 2026 AI-search study with a different corpus. It does not change these fixed measurements. AI Search Statistics for 2026 is included as context in the Sources section.
Which domains dominate Google organic results versus Perplexity Search API?
Google organic search and the Perplexity Search API produced different distributions for the same developer prompts.
Source category | Google Serper (N=250) | Perplexity API (N=250) | Total pool (N=500) |
|---|---|---|---|
Independent blogs & publishers | 100 (40.0%) | 136 (54.4%) | 236 (47.2%) |
Vendor docs & official portals | 47 (18.8%) | 75 (30.0%) | 122 (24.4%) |
Core community (Reddit/YouTube/Medium/dev.to) | 76 (30.4%) | 4 (1.6%) | 80 (16.0%) |
GitHub (repos & docs) | 13 (5.2%) | 22 (8.8%) | 35 (7.0%) |
Social & forums | 13 (5.2%) | 5 (2.0%) | 18 (3.6%) |
Academic & preprints | 1 (0.4%) | 8 (3.2%) | 9 (1.8%) |
The community-platform counts were:
Platform | Google organic results | Perplexity Search API results | Combined count |
|---|---|---|---|
| 44 (17.6%) | 0 (0.0%) | 44 (8.8%) |
| 18 (7.2%) | 0 (0.0%) | 18 (3.6%) |
| 9 (3.6%) | 0 (0.0%) | 9 (1.8%) |
| 5 (2.0%) | 4 (1.6%) | 9 (1.8%) |
Reddit was Google's most frequent domain, with 44 of 250 rows. Perplexity returned no Reddit, YouTube, or Medium rows. github.com was Perplexity's most frequent single domain, with 20 rows, or 8.0%. The broader GitHub category has 22 rows, or 8.8%, because it also includes docs.github.com. arxiv.org appeared 7 times in Perplexity rows, or 2.8%.
How does retrieval shift across different question kinds?
Retrieval mixes changed by query intent across the six question archetypes in the dataset:
Question archetype | Google community share | Perplexity community share | Perplexity vendor + GitHub share |
|---|---|---|---|
| 40.0% (18/45) | 2.2% (1/45) | 17.8% (8/45) |
| 40.0% (18/45) | 2.2% (1/45) | 0.0% (0/45) |
| 30.0% (12/40) | 0.0% (0/40) | 80.0% (32/40) |
| 27.5% (11/40) | 2.5% (1/40) | 55.0% (22/40) |
| 25.0% (10/40) | 2.5% (1/40) | 45.0% (18/40) |
| 17.5% (7/40) | 0.0% (0/40) | 42.5% (17/40) |
For comparisons, 43 of 45 Perplexity rows, or 95.6%, were in the independent category. One row came from LinkedIn and one from dev.to, so the broader remainder after community, vendor, and GitHub rows was 44 of 45, or 97.8%. For what-changed queries, Perplexity assigned 25 of 40 rows to vendor sources, or 62.5%, and 7 to GitHub, or 17.5%. Together, those categories held 80.0% of the rows. Google gave 30.0% of its version-query rows to community platforms.
Where did search candidate pools return the weakest results?
Three questions exposed clear topic mismatches in the candidate pools:
- Aider concurrent Git conflicts (Q30): Perplexity's five rows were general Git conflict documentation or GitHub pages, not Aider-specific pages. Google included Aider issue #800 alongside Stack Overflow, Reddit, and a GitHub community discussion.
- Aider versus Windsurf multi-file refactoring (Q16): No result title described a direct empirical Aider-versus-Windsurf diff benchmark. Perplexity's first result was
joinleland.com; Google mixed Aider, Cursor, and Windsurf comparisons with Reddit and dev.to pages. - Gemini 2.0 Flash tool latency (Q25): Perplexity included two iTechGuides pages, including one whose title referred to a 2026 shutdown. Google returned a Gemini API changelog and newer Gemini pages. The pool did not produce a focused latency measurement.
What does the zero for insidetheloop.dev show?
It shows that none of the 500 candidate rows used insidetheloop.dev. The CSV cannot establish whether that absence came from crawling, ranking, query wording, or another part of retrieval.
The related posts llms.txt and llms-full.txt: what they are, who reads them, and how to generate them and Cloudflare Content Signals in robots.txt: what search, ai-input, and ai-train mean document the site's crawler-facing controls. This benchmark does not test whether those controls change rankings. How AI agents fetch web pages: user agents, IP ranges, and fetch origins covers the separate fetch step.
The retrieved independent pages also show useful patterns to test, but they do not prove that any one pattern caused a ranking:
- Frontman's Roo Code versus Cline comparison starts with a direct verdict about which tool is the safer default.
- Claudefa.st's Claude Code changelog puts version-by-version release notes on one page.
- Aider's repository-map article explains tree-sitter parsing, graph ranking, and token-budget selection.
Sources
- 02-citations.csv (read 2026-10-07)
- Perplexity Search API Quickstart (read 2026-10-07)
- Serper Google Search API (read 2026-10-07)
- AI Search Statistics for 2026 (read 2026-10-07)
- Roo Code vs Cline 2026 (read 2026-10-07)
- Claude Code Changelog (read 2026-10-07)
- Building a better repository map with tree sitter (read 2026-10-07)
- llms.txt and llms-full.txt explained (read 2026-10-07)
- Cloudflare Content Signals in robots.txt (read 2026-10-07)
- How AI agents fetch web pages (read 2026-10-07)
Last verified: 2026-10-07.