Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up

At a Glance

Coordination among many language-model agents naturally concentrates effort in a small set of 'intellectual elites' following a heavy-tailed pattern, and a simple, local merge-triggering rule can rebalance that flow and improve task success by up to ~12%.

What They Found

Coordination events across large agent societies are uneven: most reasoning stays small while a few cascades grow very large, following a truncated heavy-tailed distribution (an intermediate power-law regime that cuts off due to finite resources). Early engagement on a claim increases its chance of attracting more work, producing elite agents who shoulder disproportionate cognitive effort as system size grows. Expansion actions (delegation and contradiction) scale strongly but integration (merge) does not, creating an integration bottleneck that correlates with performance drops. A focused intervention that senses when expansion outpaces integration and prioritizes merges preserves large-scale exploration but reduces wasted late-stage expansion and raises task success, especially in the most imbalanced settings. Planning Pattern

By the Numbers

11.5 million coordination events analyzed across agent counts N = {8,16,32,64,128,256,512} spanning multiple tasks and topologies.
2Coordination-event sizes fit truncated power-law behavior with estimated tail exponents α in (2, 3) across conditions, consistently favored over log-normal or pure power-law fits (p < 0.05).
3Deficit-Triggered Integration (DTI) yields relative task success improvements from +2.07% up to +12.34%, with the largest gains in planning tasks on dense mesh/fully connected topologies.

What This Means

Engineers building multi-agent systems should track how coordination is routed (not just final accuracy) because elite concentration and poor integration can cause scaling failures. Technical leaders evaluating agent reliability or multi-agent trust should add structural metrics (e.g., tail exponent, merge conversion ratio) to their dashboards to catch fragile coordination regimes before deployment. Consensus-Based Decision Pattern
Not sure where to start?Get personalized recommendations
Learn More

Key Figures

Figure 1: Heavy-tailed coordination cascades across observables. CCDFs show a power-law regime ( 2 &lt; α ^ &lt; 3 2&lt;\hat{\alpha}&lt;3 ) with truncation at large x x . Dashed lines indicate MLE fits above x min x_{\min} . Truncated power laws are favored over log-normal and exponential alternatives (Table 2 ).
Fig 1: Figure 1: Heavy-tailed coordination cascades across observables. CCDFs show a power-law regime ( 2 &lt; α ^ &lt; 3 2&lt;\hat{\alpha}&lt;3 ) with truncation at large x x . Dashed lines indicate MLE fits above x min x_{\min} . Truncated power laws are favored over log-normal and exponential alternatives (Table 2 ).
Figure 2: Finite-size stability of heavy-tailed coordination dynamics. (Left) Estimated tail exponents α ^ \hat{\alpha} (MLE) vs. agent count N N . Estimates fluctuate at small N N due to limited tail samples, then stabilize and converge beyond N ≈ 64 N\approx 64 , indicating emergence of a consistent heavy-tailed regime. (Right) Mean maximum event size ⟨ x max ⟩ \langle x_{\max}\rangle vs. N N . The upper tail grows systematically across observables, with strongest expansion for TCE, showing that increasing system size expands the reachable coordination tail.
Fig 2: Figure 2: Finite-size stability of heavy-tailed coordination dynamics. (Left) Estimated tail exponents α ^ \hat{\alpha} (MLE) vs. agent count N N . Estimates fluctuate at small N N due to limited tail samples, then stabilize and converge beyond N ≈ 64 N\approx 64 , indicating emergence of a consistent heavy-tailed regime. (Right) Mean maximum event size ⟨ x max ⟩ \langle x_{\max}\rangle vs. N N . The upper tail grows systematically across observables, with strongest expansion for TCE, showing that increasing system size expands the reachable coordination tail.
Figure 3: Topology- and task-specific heavy-tailed coordination cascades in multi-agent LLM systems. Complementary cumulative distribution functions (CCDF) of coordination-event sizes P ​ ( X ≥ x ) P(X\geq x) across four coordination observables: Delegation Cascade , Revision Wave , Contradiction Burst , and Total Coordination Effort (TCE) under four agent interaction topologies ( Chain , Star , Hierarchical , and Dynamic Reputation ) and four task families ( Reasoning , Coding , QA , and Coordination ). Each curve represents the empirical distribution of coordination-event sizes for a given task family within a topology. Across all settings, the distributions exhibit broad heavy-tailed behavior with estimated scaling exponents 2 &lt; α ^ &lt; 3 2&lt;\hat{\alpha}&lt;3 in the intermediate regime (values shown per panel), consistent with scale-free coordination dynamics. While the precise tail exponent varies modestly across tasks and architectures, the heavy-tailed form persists across all topologies, indicating that complex coordination in LLM multi-agent systems produces heterogeneous cascades spanning multiple scales. Deviations at the largest event sizes reflect finite-size truncation due to system constraints such as limited agent attention, bounded communication bandwidth, and task decomposition depth.
Fig 3: Figure 3: Topology- and task-specific heavy-tailed coordination cascades in multi-agent LLM systems. Complementary cumulative distribution functions (CCDF) of coordination-event sizes P ​ ( X ≥ x ) P(X\geq x) across four coordination observables: Delegation Cascade , Revision Wave , Contradiction Burst , and Total Coordination Effort (TCE) under four agent interaction topologies ( Chain , Star , Hierarchical , and Dynamic Reputation ) and four task families ( Reasoning , Coding , QA , and Coordination ). Each curve represents the empirical distribution of coordination-event sizes for a given task family within a topology. Across all settings, the distributions exhibit broad heavy-tailed behavior with estimated scaling exponents 2 &lt; α ^ &lt; 3 2&lt;\hat{\alpha}&lt;3 in the intermediate regime (values shown per panel), consistent with scale-free coordination dynamics. While the precise tail exponent varies modestly across tasks and architectures, the heavy-tailed form persists across all topologies, indicating that complex coordination in LLM multi-agent systems produces heterogeneous cascades spanning multiple scales. Deviations at the largest event sizes reflect finite-size truncation due to system constraints such as limited agent attention, bounded communication bandwidth, and task decomposition depth.
Figure 4: Scale-dependent emergence of broad elite tiers. (a) (a) Top- k k active agents ( E 10 active E^{\mathrm{active}}_{10} , E 25 active E^{\mathrm{active}}_{25} , E 50 active E^{\mathrm{active}}_{50} ) capture disproportionate shares of coordination effort relative to egalitarian baselines; E 10 all E^{\mathrm{all}}_{10} (dashed) confirms the result is not driven by inactive agents. of coordination effort relative to egalitarian baselines. (b) Excess concentration Δ k active \Delta^{\mathrm{active}}_{k} above equal participation increases with N N , with strongest gains in the top decile and quartile. (c) Cumulative concentration curves vs N N increasingly bow above the equality line, indicating broader and more dominant elite tiers at scale.
Fig 4: Figure 4: Scale-dependent emergence of broad elite tiers. (a) (a) Top- k k active agents ( E 10 active E^{\mathrm{active}}_{10} , E 25 active E^{\mathrm{active}}_{25} , E 50 active E^{\mathrm{active}}_{50} ) capture disproportionate shares of coordination effort relative to egalitarian baselines; E 10 all E^{\mathrm{all}}_{10} (dashed) confirms the result is not driven by inactive agents. of coordination effort relative to egalitarian baselines. (b) Excess concentration Δ k active \Delta^{\mathrm{active}}_{k} above equal participation increases with N N , with strongest gains in the top decile and quartile. (c) Cumulative concentration curves vs N N increasingly bow above the equality line, indicating broader and more dominant elite tiers at scale.

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Considerations

Experiments use a controlled set of coordination primitives and selected benchmark workloads, so dynamics might shift with different primitives or much longer-horizon interactions. Coordination was measured as discrete events (delegation, revision, contradiction, merge, total effort), which captures structural flow but can miss fine-grained semantic quality. DTI is a minimal, local intervention that improves many conditions but is not a complete solution—dynamic topology changes and longer-term mechanisms for distributing integration remain open directions. Context Drift

Methodology & More

Runs across four agent benchmarks (QA, reasoning, coding, planning), seven society sizes (8–512 agents), and multiple interaction topologies logged individual reasoning actions and the claim-to-claim cascades they produced. Measuring five atomic coordination events—delegation, revision, contradiction, merge, and total cognitive effort—revealed that most trajectories are small while a minority become very large. Those large cascades follow a truncated heavy-tailed (power-law) form with tail exponents between 2 and 3, indicating a scale-free amplification over an intermediate range that is cut off by finite resources like context and communication limits. Micro-level reinforcement explains the macro pattern: claims that receive early engagement tend to attract more follow-up (a preferential-attachment effect), producing emergent intellectual elites whose share of effort grows with society size. Expansion mechanisms (delegation and contradiction) drive cascade growth, but merge operations that consolidate branches do not scale at the same rate, creating an integration bottleneck that correlates with performance collapse in high-intensity regimes. Monitoring the balance between expansion and integration and applying Deficit-Triggered Integration—a local rule that temporarily prioritizes merges when imbalance exceeds a threshold—keeps large-scale exploration intact while converting late-stage expansion into earlier integration. That change preserves the heavy-tailed structure but reduces inefficient tail mass and improves task success (up to ~12%), especially in deep planning tasks on dense topologies. The findings argue that reliable scaling of multi-agent setups requires design and monitoring of coordination structure, not only stronger individual agents. Guardrails Pattern
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

ArXiv preprint with no listed affiliations and low h-indexes (4 and 5). Limited author reputation signals -> emerging/limited info.