Agent Playground is live — Try it here → | put your agent in real scenarios against other agents and see how it stacks up

The Big Picture

Two bits per message, when interpreted by each receiver using its own local context, preserve the decision-relevant differences receivers need and often match or outperform methods that send thousands of bits.

The Evidence

Messages should preserve the receiver’s differences in action value within the receiver’s own situation, not an average across all situations. Implementing that idea gives two practical estimators: an offline exact codebook (teacher builds the mapping, then sender is frozen) and an online teacher-free method (a paired reference branch aligns discrete symbols to continuous sender states). Across navigation, predator–prey, StarCraft and multi-agent particle tasks, the conditional, receiver-aware messages delivered large gains in return and coordination while using only 2 bits per message. receiver-conditioned messaging Across navigation, predator–prey, StarCraft and multi-agent particle tasks, the conditional, receiver-aware messages delivered large gains in return and coordination while using only 2 bits per message. estimators
Not sure where to start?Get personalized recommendations
Learn More

Data Highlights

1RAVEN won 38 of 40 seed-paired comparisons against strong baselines on navigation tasks.
2The offline method required 82–98% fewer execution floating-point operations than several competitors (NDQ, CACOM, ExpoComm).
3Online RAVEN’s 2-bit channel increased predator–prey capture success by 42.8 percentage points.

What This Means

Engineers building agents that must coordinate over extremely tight bandwidth (radio, acoustic, or protocol-limited links) can use the method to get strong coordination with tiny messages. Technical leads evaluating multi-agent stacks should consider receiver-conditioned messaging to reduce communication cost and improve robustness. Researchers in multi-agent communication will find the conditional target and the offline/online estimators a practical way to study decision-focused compression. multi-agent data analysis

Key Figures

Figure 18. Final score of each method on each task, colored by the per-task min–max normalized score (cell values are win rates or returns; the mean normalized score is in brackets).
Fig 18: Figure 18. Final score of each method on each task, colored by the per-task min–max normalized score (cell values are win rates or returns; the mean normalized score is in brackets).

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Yes, But...

RAVEN helps most where a few symbols must carry information the receiver cannot observe and where the receiver’s context changes message meaning; in larger multi-hop rings the benefit can shrink. The offline estimator needs a centralized teacher and an enumerable joint action space to build the exact codebook. The online estimator removes the teacher but has more parameters (24–54% larger in some benchmarks) and can learn more slowly early on the hardest maps. centralized teacher conditioning on the receiver

Methodology & More

Agents that share a tiny fixed alphabet (four symbols = two bits) must pick which distinctions to keep. The key idea is to measure what matters to the receiver in its own situation: preserve the receiver’s centered action-value differences conditioned on a small summary of the receiver’s context. One symbol can mean different actions in different receiver situations because each receiver decodes the symbol using its private context. Two implementations realize that principle. Offline RAVEN uses a centralized teacher to tabulate receiver-conditioned action values, enumerates the optimal 4-symbol codebook exactly, distills it into a sender network and then freezes the sender while training receivers. Online RAVEN removes the teacher by adding a training-only reference branch inside a value-factorization learner that replaces decoded symbols with the sender’s continuous state; an alignment loss pulls the deployed discrete branch toward the reference. Empirically, RAVEN consistently improves returns across navigation tasks and multi-agent benchmarks, wins the majority of seed-paired comparisons, and achieves large gains (for example, +42.8 percentage points in predator–prey) while using only two bits per message. Analysis shows conditioning on the receiver explains most of the gain, and the conditional reconstruction risk correlates with actual returns, supporting the proposed mechanism. Practical trade-offs: offline RAVEN is strongest when a teacher is available; online RAVEN is broadly applicable but costs more parameters and can train slower on some hard problems. Orchestration
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

One author has low h-index (8) and no affiliations or venue prestige; arXiv preprint with no citations indicates emerging/limited information.