At a Glance
Top large language models working as teams complete only about half of long, multi-step tasks, and fewer than one in three agent actions actually help reach the goal.
ON THIS PAGE
What They Found
AgentWorld is a long-horizon simulator built to test how groups of AI agents coordinate over 25–55 rounds with 3–20 specialized roles. Teams using current large language models often start tasks but fail at late-stage coordination: best success rate is 52.0% and causal analysis shows most actions do not contribute to outcomes. Common failure modes include poor information sharing, role confusion (duplicate or off-role work), and inability to keep a shared plan across many rounds. The benchmark and a new causal metric are open-sourced so teams can test and improve multi-agent coordination. role-based coordination pattern
Not sure where to start?Get personalized recommendations
Data Highlights
1Best model success rate: 52.0% task success (top of the four evaluated models).
2Causal Collaboration Effectiveness (CCE) for best model: 0.320 — under one-third of agent actions causally contributed to success.
327% of tasks were unsolved by any model; example gap between partial and full progress: 71.5% partial success rate vs. 52.0% full success for one evaluated model.
What This Means
Engineers building multi-agent systems should use this to spot coordination bottlenecks before deployment and to focus work on communication, role assignment, and persistent planning. Technical leaders evaluating agent teams can use these results to set realistic expectations about reliability and to prioritize monitoring and trust signals. Researchers can adopt the open benchmark and causal metric to compare solutions and study failure modes systematically. trust signals
Key Figures

Fig 1: Figure 1: An example of the AgentWorld environment with 10 agents interacting through live chat.

Fig 2: Figure 2: Overview of AgentWorld. The MMORPG sandbox with diverse biomes and resources (left), blackbox multi-agent interaction model (center), and 10 representative high-level API tools (of 13 total) that abstract away low-level game mechanics (right).

Fig 3: Figure 3: Example task definition (Task 86: Forge Vanguard). Ten agents collaborate through a mine → \rightarrow smelt → \rightarrow craft pipeline: miners extract ore and coal, the smelter processes them into iron bars, and smiths forge heavy swords and an axe. Each agent has unique skills, equipment, inventory, and spawn location.

Fig 6: Figure 6: Task annotation platform. (a) World map for selecting agent spawn locations. (b) Skill level configuration with sliders. (c) Searchable inventory item browser (133 items). (d) Equipment slots with character preview.
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreYes, But...
The environment is a simulated MMORPG that abstracts low-level actions into high-level tools, so results reflect collaboration quality more than fine motor control and may not map perfectly to every real-world domain. Only four frontier models were tested under a fixed prompt setup—different models, prompts, or training could change outcomes. The causal metric reduces subjectivity by using models to build action graphs, but it still depends on automated judgments and may miss subtle, indirect contributions. Tree of Thoughts Pattern
Methodology & More
AgentWorld is a purpose-built sandbox for long, collaborative missions where 3–20 specialized agents must coordinate over 25–55 rounds to complete tasks like supply chains, crafting pipelines, and large-scale resource transfers. Low-level mechanics (movement, combat, crafting details) are hidden behind 13 high-level API tools so evaluations focus on planning, messaging, and role coordination rather than action execution. The benchmark offers 100 human-annotated tasks plus 100 variants, and a novel metric called Causal Collaboration Effectiveness (CCE) that traces which agent actions actually caused progress toward the goal by building causal action graphs. Four leading large language models were run on the benchmark with identical prompts and tool definitions to isolate collaboration ability. Results show a clear gap: the best model solved only 52.0% of tasks and had a CCE of 0.320, meaning most agent actions were irrelevant to success. Failures cluster around breakdowns in communication (missing or late information), role confusion (redundant or off-role behavior), and plan drift over many rounds. The open sandbox and metric let teams reproduce these findings, track agent-to-agent evaluation, and iterate on fixes like better explicit role assignment, shared plan storage, and trust/reputation signals to reduce wasted actions and improve reliability. Mutual verification pattern
Avoid common pitfallsLearn what failures to watch for
Credibility Assessment:
Contains a well-established author (Lyle Ungar, h-index 21) indicating an established researcher; despite arXiv venue and mixed h-indexes, presence of a senior researcher raises credibility to 4.