Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up

At a Glance

Don't trust the common 'rejection rate' metric — it can be arbitrarily manipulated and does not predict real error. Measure agents by their actual error probability, because decentralization adds an irreducible gap set by the network and by initialization/prior choices.

What They Found

Rejection rate (the long-run average log-belief ratio) can be increased arbitrarily without improving, and sometimes while worsening, actual decision error — so it is a poor comparison metric. For binary Gaussian detection, the long-run ratio between a decentralized agent's error and the optimal centralized error factors into two multiplicative penalties: one due to imperfect network connectivity and one due to initial beliefs or wrong priors. Even when decentralized schemes match the best possible exponential decay (same error exponent), these higher-order factors produce an irreducible performance gap that varies by agent and network position. irreducible performance gap
Not sure where to start?Get personalized recommendations
Learn More

By the Numbers

1Any rejection-rate value is achievable: given any algorithm and any factor b>1, you can construct another scheme whose rejection rate is multiplied by b (so the rejection rate can be made arbitrarily larger).
2For the Gaussian binary shift-in-mean case, the asymptotic error ratio includes α, where α = 1/K for traditional geometric-averaging pooling and α = 1 for the corrected weighting scheme — so weighting choice changes the non-exponent term by a factor scaling with K.
3Empirical checks used large Monte Carlo trials (10^3–10^4 runs) to confirm theory: networks that are not fully connected show multiplicative error gaps exp(ε_k)>1, and nonuniform initial beliefs add a further multiplicative penalty (visible in the plots comparing central vs. peripheral agents).

What This Means

Engineers building cooperative multi-agent systems who need reliable decision outcomes — use error probability, not rejection rate, to compare designs. Technical leads deciding network topology or pooling weights — the network design and initial belief/prior setup create an unavoidable performance gap that affects some agents more than others. Researchers developing agent evaluation tools (agent-to-agent evaluation) and trust metrics — avoid using rejection-rate style signals as stand-ins for actual reliability.

Key Figures

Figure 1: Rejection rate vs. error probability. Traditional SL behavior in the setting of Example 1 , for the network topology shown in the inset panel of Fig. 2 . ( Left ). One realization of the belief 𝝁 1 , t ​ ( θ ∙ ) \bm{\mu}_{1,t}(\theta^{\bullet}) where, after a random time instant 𝒕 0 ≈ 25 \bm{t}_{0}\approx 25 , the beliefs of the compared strategies follow the ordering dictated by the rejection rates. ( Middle ). A different realization of the belief 𝝁 1 , t ​ ( θ ∙ ) \bm{\mu}_{1,t}(\theta^{\bullet}) , where the strategies follow the ordering of the rejection rates after a longer time, i.e., 𝒕 0 ≈ 130 \bm{t}_{0}\approx 130 . Note that the modified quantized implementation is characterized by a more significant variability of the beliefs with respect to traditional SL. ( Right ). Error probability curves obtained by averaging over multiple realizations. The quantized implementations share the same error probability curves in view of ( 24 ), but they feature a higher error probability compared to traditional SL.
Fig 1: Figure 1: Rejection rate vs. error probability. Traditional SL behavior in the setting of Example 1 , for the network topology shown in the inset panel of Fig. 2 . ( Left ). One realization of the belief 𝝁 1 , t ​ ( θ ∙ ) \bm{\mu}_{1,t}(\theta^{\bullet}) where, after a random time instant 𝒕 0 ≈ 25 \bm{t}_{0}\approx 25 , the beliefs of the compared strategies follow the ordering dictated by the rejection rates. ( Middle ). A different realization of the belief 𝝁 1 , t ​ ( θ ∙ ) \bm{\mu}_{1,t}(\theta^{\bullet}) , where the strategies follow the ordering of the rejection rates after a longer time, i.e., 𝒕 0 ≈ 130 \bm{t}_{0}\approx 130 . Note that the modified quantized implementation is characterized by a more significant variability of the beliefs with respect to traditional SL. ( Right ). Error probability curves obtained by averaging over multiple realizations. The quantized implementations share the same error probability curves in view of ( 24 ), but they feature a higher error probability compared to traditional SL.
Figure 2: Network error. Error probabilities of the considered SL strategies for the decision-making problem in Sec. VI-A . In this experiment, the prior and the initial beliefs are uniform. For the decentralized implementations, the agents are connected according to the network topology shown in the inset panel (here the edges are depicted with lines and no arrows, since they are undirected ). The error probabilities of agents 2 2 and 3 3 are shown, since these two agents represent a central and a peripheral agent. This choice highlights the effect of the network topology on the performance. The arrows emphasize the network error exp ⁡ ( ε k ) \exp(\varepsilon_{k}) in ( 29 ). In agreement with Theorem 1 , traditional SL and NB 2 feature the same curves because of the uniform initialization. All the curves are obtained averaging over 10 4 10^{4} Montecarlo runs.
Fig 2: Figure 2: Network error. Error probabilities of the considered SL strategies for the decision-making problem in Sec. VI-A . In this experiment, the prior and the initial beliefs are uniform. For the decentralized implementations, the agents are connected according to the network topology shown in the inset panel (here the edges are depicted with lines and no arrows, since they are undirected ). The error probabilities of agents 2 2 and 3 3 are shown, since these two agents represent a central and a peripheral agent. This choice highlights the effect of the network topology on the performance. The arrows emphasize the network error exp ⁡ ( ε k ) \exp(\varepsilon_{k}) in ( 29 ). In agreement with Theorem 1 , traditional SL and NB 2 feature the same curves because of the uniform initialization. All the curves are obtained averaging over 10 4 10^{4} Montecarlo runs.
Figure 3: Belief initialization. Same setup used in Fig. 2 , but for the prior and the initial beliefs, which are now set to [ 0.65 , 0.35 ] [0.65,0.35] . Comparing against Fig. 2 , and in agreement with Theorem 1 , we see that now the error probability for traditional SL pays an additional error term quantified (in the logarithmic scale adopted in the figure) by log 10 ⁡ cosh ⁡ ( ( K − 1 ) ​ ξ / 2 ) \log_{10}\cosh(\,(K-1)\,\xi/2\,) , due to the non-uniform initialization, yielding an overall gap equal to ε k ​ log 10 ⁡ e + log 10 ⁡ cosh ⁡ ( ( K − 1 ) ​ ξ / 2 ) \varepsilon_{k}\,\log_{10}e+\log_{10}\cosh(\,(K-1)\,\xi/2\,) . The arrows in the figure emphasize the two additive terms of this gap.
Fig 3: Figure 3: Belief initialization. Same setup used in Fig. 2 , but for the prior and the initial beliefs, which are now set to [ 0.65 , 0.35 ] [0.65,0.35] . Comparing against Fig. 2 , and in agreement with Theorem 1 , we see that now the error probability for traditional SL pays an additional error term quantified (in the logarithmic scale adopted in the figure) by log 10 ⁡ cosh ⁡ ( ( K − 1 ) ​ ξ / 2 ) \log_{10}\cosh(\,(K-1)\,\xi/2\,) , due to the non-uniform initialization, yielding an overall gap equal to ε k ​ log 10 ⁡ e + log 10 ⁡ cosh ⁡ ( ( K − 1 ) ​ ξ / 2 ) \varepsilon_{k}\,\log_{10}e+\log_{10}\cosh(\,(K-1)\,\xi/2\,) . The arrows in the figure emphasize the two additive terms of this gap.
Figure 4: Lack of prior knowledge. Same setup used in Fig. 2 , but for the true prior, which is now set to [ 0.65 , 0.35 ] [0.65,0.35] . However, the agents do not know the true prior and continue to use uniform initial beliefs as in Fig. 2 .
Fig 4: Figure 4: Lack of prior knowledge. Same setup used in Fig. 2 , but for the true prior, which is now set to [ 0.65 , 0.35 ] [0.65,0.35] . However, the agents do not know the true prior and continue to use uniform initial beliefs as in Fig. 2 .

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Yes, But...

The exact closed-form gap is derived only for the binary Gaussian (shift-in-mean) model under independence across agents and time, so quantitative constants may change for correlated or non-Gaussian data. Results assume primitive left-stochastic or doubly-stochastic pooling matrices; pathological graph or communication models (e.g., time-varying links or strong dependencies) require fresh analysis. While the error exponent (the exponential decay rate) can match the centralized optimum, higher-order terms still matter in practical finite-time regimes and can dominate observed error behavior. time-varying links

Methodology & More

Rejection rate (the limiting average log-belief ratio) is commonly used to compare decentralized learning strategies, but it can be arbitrarily inflated without improving true decision quality. Constructive examples show paradoxes: you can make a 1-bit-quantized scheme appear better by rejection-rate while it actually loses information and has higher error; you can discard most agents' data yet increase rejection rate; and the exact Bayesian posterior can be outperformed in rejection-rate while being optimal in error probability. These contradictions arise because rejection rate is a limit-of-averages statistic and ignores the random variability and finite-time mistakes that determine error probability. fully connected The network term equals 1 only for fully connected (ideal) mixing; otherwise it strictly increases the error. The initialization/prior term can be neutralized by using corrected weights (their “NB2”-style weighting yields α=1, while traditional geometric pooling gives α=1/K), but the network contribution remains irreducible. Numerical experiments (averaged over thousands of Monte Carlo runs) confirm that different agents experience persistent, different gaps depending on their graph position, and that rejection rate is a misleading metric for practical evaluation. limit-of-averages
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

Low reported h-indexes and no prominent affiliations or venue; limited signals of established credibility.