Microsoft Research's Agensh is a multi-agent harness that ran up to 1,024 identical coding agents without a central lead. In the paper submitted on 2026-09-22, one agent passed 33.89% of pandoc's hidden tests and 1,024 agents passed 55.06% under a six-hour, no-Internet budget.
Agensh's key facts
- The arXiv record lists the paper "Agensh: Scaling Organizational Intelligence to 1,024 Agents" and records its first submission on 2026-09-22.
- Every Agensh worker runs the same asynchronous cooperation loop. Workers gather context, claim and self-assign sub-tasks, act, share findings, verify results, and merge progress.
- The evaluation used GPT-5.6-sol (high) on the five hardest ProgramBench tasks: FFmpeg, gromacs, pandoc, PHP-src, and ctags.
- Across those five tasks, the mean final test-pass rate rose from 19.31% with 1 agent to 28.78% with 128 agents.
- On pandoc, the final test-pass rate rose from 33.89% with 1 agent to 55.06% with 1,024 agents.
- The paper reports no total token consumption, dollar cost, or server compute cost for the runs.
How Agensh coordinates agents without a lead
Agensh replaces a lead session with shared state. Each worker repeats the same five-step loop:
- Gather context. Read the goal, recent repository work, messages, and shared findings.
- Claim a sub-task. Pick an unaddressed area, record the claim, and message peers when work overlaps.
- Take action. Work in a local container and publish useful findings as they become available.
- Verify results. Compare the implementation with reference behavior under the benchmark's local limits.
- Merge progress. Push the change, open and resolve a pull request, merge it, and record a summary.
Three shared components carry the coordination:
- Shared workspace: Gitea stores the repository, branches, issues, and pull requests.
- Message interface: Mattermost carries team announcements and direct messages between workers.
- Shared context: An append-only board stores typed findings such as
OBSERVED,FACT,FAIL,CLAIM, andPATCH_SUMMARY. Workers can search older entries instead of relying only on their current context.
The paper reports different coordination patterns as the team grows. Eight agents negotiate interfaces and avoid duplicate work. Thirty-two agents add peer review. At 128 agents, workers standardize integration and pull-request checks. At 1,024 agents, several workers act as integrators and peers can take over when an integrator does not respond.
Agensh's ProgramBench results
ProgramBench asks agents to rebuild software from scratch and checks the result with hidden behavioral tests. Agensh used a six-hour budget without Internet access. The five-task mean was measured through 128 agents. The 1,024-agent result was reported for pandoc, not as a five-task mean.
Worker count | Mean final pass rate across five tasks | pandoc final pass rate |
|---|---|---|
1 agent | 19.31% | 33.89% |
8 agents | 20.68% | Not reported |
32 agents | 26.52% | Not reported |
128 agents | 28.78% | 50.94% |
1,024 agents | Not reported | 55.06% |
The larger teams also reached an intermediate score sooner on pandoc. The 128-agent run passed 30% of tests at the 30-minute checkpoint. The 32-agent run reached that mark at 60 minutes, and the 8-agent run at 90 minutes. The single-agent run stayed below 30% during the first two hours.
The result measures benchmark pass rates, not the cost of obtaining them. The paper does not report total tokens, per-run spend, or server compute expenses. That omission prevents a cost-per-gain calculation from the paper alone.
Agensh's code link and project page
The project page linked by Microsoft's aka.ms/Agensh shortlink is available. On 2026-10-06, the shortlink returned an HTTP 301 redirect to the project page:
HTTP/1.1 301 Moved Permanently
Location: https://agens-harness.github.io/project/
HTTP/2 200The paper also cites https://github.com/microsoft/Agensh. On 2026-10-06, that URL returned HTTP 404:
HTTP/2 404The cited GitHub URL therefore did not provide a cloneable repository on 2026-10-06. That does not establish whether the source code exists somewhere else.
Sources
- Agensh: Scaling Organizational Intelligence to 1,024 Agents: https://arxiv.org/abs/2609.26781 (read 2026-10-06)
- Agensh project page: https://agens-harness.github.io/project/ (read 2026-10-06)
- Microsoft Agensh shortlink: https://aka.ms/Agensh (read 2026-10-06)
- Microsoft Agensh code URL: https://github.com/microsoft/Agensh (read 2026-10-06)
- ProgramBench extended results: https://programbench.com/extended/ (read 2026-10-06)
Last verified: 2026-10-06.