Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up

At a Glance

Large language models tend to default to the same actions, and while they can shift when coordination rewards alignment, they struggle to remain different when diversity is actually more valuable.

What They Found

Agents built from large language models show a strong baseline tendency to pick the same actions as one another. They do adjust that similarity when incentives clearly reward matching, achieving very high coordination when aligned. However, when the game rewards doing different things, humans sustain diverse strategies more often than these models. That gap suggests models may create fragile group behavior unless designers explicitly reward or enforce useful differences. Emergence-Aware Monitoring Pattern
Not sure where to start?Get personalized recommendations
Learn More

Data Highlights

1LLMs showed high baseline action similarity in roughly 75% of rounds — about 1.5× higher than human subjects (~50%).
2When payoffs favored matching actions, LLM groups achieved near-perfect coordination (~95% success), slightly above human groups (~90%).
3When divergence was rewarded, humans maintained heterogeneous strategies about 65% of the time, while LLMs did so in only ~35% of rounds — a ~30 percentage-point gap.

What This Means

Engineers building systems where multiple agents interact should care because default model behavior can cause over-similarity that breaks flexible coordination. Product leaders and evaluators running agent-to-agent evaluation or monitoring multi-agent trust need to test for both the models' tendency to match and their ability to stay different when diversity is beneficial.

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Yes, But...

The experiments used simplified coordination games; real-world tasks are often more complex and may change these numbers. Results depend on the specific language models, prompts, and incentive structures used, so behavior can vary across deployments. Human subjects and models face different costs and information — human incentives, communication, or prior experiences may explain part of the gap in sustaining diversity. Guardrails Pattern

Methodology & More

Researchers separated two ideas: baseline similarity (agents independently choosing similar actions) and strategic similarity (agents changing how similar they are in response to rewards). They ran clean coordination games with both human participants and large language model agents to observe how each group behaved when the best outcome required either matching choices or doing different things. Findings show that large language model agents naturally pick similar actions a lot of the time, which makes them excellent at coordination when everyone benefits from matching. However, when the highest payoff requires agents to be different, models were much less likely than humans to maintain the needed diversity. For practitioners, that means multi-agent deployments may work great when tasks favor consensus but can fail or become brittle when complementary or diverse actions are required. Practical fixes include explicit incentives for diversity, randomized decision components, or evaluation protocols that measure both alignment and the ability to sustain useful differences during agent-to-agent evaluation and continuous monitoring. Emergence-Aware Monitoring Pattern Continuous Monitoring
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

All authors have very low h-indices (1–4), no affiliations or venue beyond arXiv, and zero citations — fits emerging/limited-info category.