Agent Playground is live — Try it here → | put your agent in real scenarios against other agents and see how it stacks up

At a Glance

On tasks where one correct answer matters, larger teams hit a voting ceiling tied to the model’s most common answer; on numeric estimation tasks, simple averaging barely helps because item-level bias dominates, so diversity or better base models matter more than scale.

What They Found

Scaling behaves very differently depending on the task type. For multiple-choice or problem-solving tasks, sampling more answers raises the chance of finding a correct solution but plurality voting quickly converges to the model’s modal answer, creating a hard ceiling that matches a simple independence model within 0.5 percentage points. For numeric “guesstimate” tasks, most error comes from systematic, item-level bias—about 87% of squared log-error—so averaging many samples reduces error by only about 6%. Heterogeneous teams (mixing model families) improve over the average member and help on estimation tasks, but on disjunctive tasks they still rarely beat the single best model in the group. Consensus-Based Decision Pattern
Not sure where to start?Get personalized recommendations
Learn More

By the Numbers

1Item-level bias explains about 87% of the squared log-error in the Fermi (estimation) benchmark.
2Averaging multiple estimates reduces error by only ~6% on compensatory (estimation) tasks.
3Plurality voting behavior across teams up to 30 agents is predicted by a conditional-independence model to within 0.5 percentage points; revision rounds improve accuracy, but seeing 1 peer or 29 peers gives nearly the same lift.

What This Means

Engineers building multi-agent systems and technical leaders deciding whether to scale agent counts: use these results to choose aggregation strategies (vote, verify, or average) and whether to invest in diversity instead of sheer scale. Researchers and evaluation teams: expect different scaling laws for problem-solving vs estimation tasks, and design benchmarks and agent-to-agent evaluation accordingly. Mutual Verification Pattern

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Considerations

Experiments used an “answer-first” prompt style that lowers absolute accuracy compared with step-by-step reasoning, so ceilings may shift under other prompting. All models were open-weight and at most 20 billion parameters; larger or proprietary models may behave differently. The estimation results rely heavily on one benchmark that contains some noisy or implausible reference labels, which the authors note and partially control for in analysis. Tree of Thoughts Pattern

Methodology & More

They mapped common benchmarks to two group-task types: disjunctive tasks (where group success needs at least one agent to find a correct solution and the group then select it) and compensatory tasks (where the group aggregates continuous numeric estimates). Using 13 open models ranging from 3B to 20B parameters, the team ran up to 30-agent ensembles and up to three rounds of peer-aware revision. Disjunctive results were aggregated with plurality voting and pass@N-style metrics; estimation tasks used geometric mean and measured absolute log-error. Findings are clear and practical. For disjunctive problems (math, multiple-choice, code), sampling more answers helps but plurality voting rapidly converges to a model’s most frequent answer, producing a ceiling that can be predicted by a conditional-independence model. Allowing agents to revise after seeing peers helps, but one peer often gives almost the same benefit as many peers. For compensatory estimation tasks, nearly all error is item-specific bias, so naive averaging yields only marginal gains (~6% error reduction). Mixing model families (heterogeneous teams) offers the best practical win: teams beat their average member and can outperform individuals on estimation tasks, though on disjunctive tasks they still rarely surpass the strongest single model. Practical implication: don’t assume more identical agents will solve your problems — prefer verifiers, better base models, or diverse model mixes depending on task type. Role-Based Agent Pattern The paper also formalizes these limits and provides bounds that practitioners can use to predict when voting or averaging will stop improving performance. For evaluation and trust tracking, the results argue for agent-to-agent evaluation and monitoring model biases at the item level rather than relying solely on scaling agent counts. Consensus-Based Decision Pattern
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

Authors have low h-indexs (e.g., 4) and no clear strong institutional affiliations; arXiv preprint with no citations — Emerging / limited info.