The Big Picture
More cooperating workers steadily improves both speed and final accuracy: increasing worker count is a practical lever to get better results on long, complex software tasks without a central manager.
ON THIS PAGE
The Evidence
A shared workspace and simple, self-driven loop let many independent workers coordinate without a central orchestrator. As the number of workers grows, the group both reaches higher final success rates and gets to those results faster. At larger sizes the organization also develops clearer division of labor and repeatable workflows, letting specialized roles and multi-worker integrations emerge naturally.
Data Highlights
1Across five very hard software-reconstruction tasks, going from 1 to 128 workers raised the mean final test-pass rate from 19.31% to 28.78% — a 9.47 percentage-point (≈49% relative) improvement.
2On the pandoc task, final test-pass rates climbed from 33.89% with 1 worker to 50.94% with 128 workers and 55.06% with 1,024 workers (a 21.17 percentage-point absolute gain over a single worker).
3More workers cut time to milestones: on pandoc, 128 workers crossed a 30% pass-rate at 30 minutes, versus 32 workers at 60 minutes and 8 workers at 90 minutes; a single worker stayed below 30% for the first two hours.
What This Means
Platform engineers and technical leads building systems that split work across many AI agents should care because agent count is an actionable way to improve throughput. Throughput researchers and evaluation teams should consider agent count as a distinct scaling axis when comparing multi-worker setups or designing benchmarks.
Not sure where to start?Get personalized recommendations
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreConsiderations
Results used a fixed underlying model and a 6-hour, no-internet software reconstruction benchmark, so gains may depend on task type and model quality. Running hundreds or thousands of workers brings compute and cost trade-offs that the experiments do not fully quantify. Real-world failure modes (e.g., conflicting merges) may differ outside the tested code-reconstruction tasks.
Methodology & More
Agensh replaces a central planner with lightweight shared infrastructure (a workspace for merged work, a message channel for announcements and direct messages, and a shared context of reusable findings) and a simple five-step loop each worker follows: gather context, claim a sub-task, act with tools, verify results, and merge progress. Workers operate concurrently and asynchronously, claiming tasks themselves and resolving overlap via messages, so verified contributions accumulate without a single orchestrator.
The team evaluated Agensh on five extremely demanding code-reconstruction tasks from ProgramBench under a 6-hour budget with a fixed agent model. Increasing worker count consistently improved final test-pass rates and reduced time to reach performance milestones. The experiments scaled up to 1,024 workers on one task (pandoc), showing continued gains and the spontaneous emergence of standardized workflows and specialized roles as organization size increased. The takeaway: agent count is a practical, new scaling dimension for multi-worker systems, but adopting it requires weighing compute cost, task fit, and communication design when moving beyond code-focused benchmarks.
Avoid common pitfallsLearn what failures to watch for
Credibility Assessment:
ArXiv preprint but multiple authors with moderate h-index (h≈9–11) and a recognizable name (Furu Wei) associated with top industry research, indicating stronger credibility.