benchmarks
Which sites search engines surface for 50 AI-agent questions: our October 2026 measurement
For 50 AI-agent questions, Google gave community sites 30.4% of candidate rows; Perplexity gave them 1.6%, with more vendor documentation and GitHub.
AXI is ten rules for CLIs that agents run
AXI defines ten design principles for agent-run CLIs across four official tools: gh-axi, chrome-devtools-axi, lavish-axi, and quota-axi.
GitHub ReviewBench scores AI code review agents on 219 pull requests
GitHub ReviewBench compares AI code review agents on 219 pull requests with grounded and augmented precision, recall, and F1 metrics.
What a System One model decides, and what it refuses
System One models return typed probabilities for bounded questions and refuse free-form text, code, and reasoning explanations.
Clef, Clef-flash, or Jev: which decision model belongs on the agent hot path?
Use Clef-flash for speed and vision, Clef for a larger model and context window, or Jev for low-cost text routing.
Liquid's d1 now answers from an image
Liquid AI added vision to its d1 decision model on 2026-10-05, returning calibrated probabilities at $0.04 per million input tokens with zero output tokens.
Microsoft ran 1,024 coding agents with no lead
Microsoft's Agensh ran 1,024 coding agents without a lead and lifted pandoc's test-pass rate from 33.89% to 55.06%.