Tag

benchmarks

  1. 7 min

    Which sites search engines surface for 50 AI-agent questions: our October 2026 measurement

    For 50 AI-agent questions, Google gave community sites 30.4% of candidate rows; Perplexity gave them 1.6%, with more vendor documentation and GitHub.

  2. 5 min

    AXI is ten rules for CLIs that agents run

    AXI defines ten design principles for agent-run CLIs across four official tools: gh-axi, chrome-devtools-axi, lavish-axi, and quota-axi.

  3. 6 min

    GitHub ReviewBench scores AI code review agents on 219 pull requests

    GitHub ReviewBench compares AI code review agents on 219 pull requests with grounded and augmented precision, recall, and F1 metrics.

  4. 5 min

    What a System One model decides, and what it refuses

    System One models return typed probabilities for bounded questions and refuse free-form text, code, and reasoning explanations.

  5. 6 min

    Clef, Clef-flash, or Jev: which decision model belongs on the agent hot path?

    Use Clef-flash for speed and vision, Clef for a larger model and context window, or Jev for low-cost text routing.

  6. 5 min

    Liquid's d1 now answers from an image

    Liquid AI added vision to its d1 decision model on 2026-10-05, returning calibrated probabilities at $0.04 per million input tokens with zero output tokens.

  7. 4 min

    Microsoft ran 1,024 coding agents with no lead

    Microsoft's Agensh ran 1,024 coding agents without a lead and lifted pandoc's test-pass rate from 33.89% to 55.06%.