Agent Playground is live — Try it here → | put your agent in real scenarios against other agents and see how it stacks up

Key Takeaway

Simulated AI agent groups produce much higher rates of full agreement than matched human groups (by roughly 34–44 percentage points), but that agreement often ignores correctness, so consensus is a weak trust signal.

Key Findings

When replaying 100 human discussion groups on a logic task with matched AI agent groups seeded with each participant's initial belief, simulated groups reached far higher agreement than people. Consensus-Based Decision Pattern Human full-consensus rates depended a lot on scoring choices (24.0%–57.0%), while about one fifth of human participants never posted but agents nearly always contributed. Across two different ways of compensating for those measurement differences, agent groups still showed gaps of about 34 to 44 percentage points in full consensus, with reasoning-mode agents often reaching near-unanimous but incorrect answers. Simulated consensus did not predict collective accuracy, and the belief-seeded agents were biased estimators of human group outcomes.

Data Highlights

1Human full-consensus rates varied from 24.0% to 57.0% depending on how participation and final states were scored.
2About 20% of human participants never posted during group discussion, whereas seeded AI agents almost always produced a contribution.
3Simulated agent groups showed gaps of roughly 34.0 to 44.4 percentage points higher full-consensus than humans across submit-based and participation-matched comparisons (e.g., 34.0 and 43.9 points for chat and reasoning modes).

Implications

Engineers building multi-agent systems and teams that use simulated deliberation for evaluation should care because consensus among agents can overstate real-world agreement and mask errors. Technical leads and researchers using agent-to-agent evaluation or reputation signals should test participation patterns and accuracy, not rely on unanimous answers as a sign of trustworthiness. Semantic Capability Matching Pattern
Explore evaluation patternsSee how to apply these findings
Learn More

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Limitations

The study replayed one classic reasoning task (the Wason selection task) with 100 held-out human groups, so results may differ for other tasks, group sizes, or domain knowledge. Agents were seeded with each participant's pre-discussion answer; different seeding strategies or agent architectures could change outcomes. Scoring and participation definitions strongly affect measured consensus, and while two complementary sensitivity analyses converged here, other measurement choices could produce different gaps. Inter-Agent Miscommunication

Full Analysis

Researchers recreated 100 human discussion groups on a logic puzzle by spawning matched AI agent groups: each agent inherited a single participant's pre-discussion belief and then interacted under the same scoring code used for the humans. Two main agent modes were tested (a chat mode and a reasoning mode), and several scoring approaches were applied to make human and agent behavior comparable. To address measurement asymmetries—like humans who never posted—the study ran two sensitivity analyses: one limited to people who submitted answers and another that matched participation patterns between people and agents. They also tested removing a memorized correct answer and disabling early stopping. Orchestrator-Worker Pattern Hallucination Propagation
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

Single author, no affiliations or h-index provided, arXiv preprint with zero citations — lacks recognizable institutional or author reputation.