Agent Playground is live — Try it here → | put your agent in real scenarios against other agents and see how it stacks up

Key Takeaway

You can reuse an already-trained single-agent controller unchanged and add a small input adapter to handle teammates, updating only 3.5–7.3% of parameters while matching or beating multi-agent training from scratch.

What They Found

A compact input-side adapter translates multi-agent observations into the format expected by a frozen single-agent policy, letting the original controller drive actions without retraining. This aligns with the Semantic Capability Matching Pattern. Training only the adapter (and optional small critic head) yields better performance than training agents from scratch across three diverse tasks. The approach is highly parameter-efficient and generalizes to team sizes not seen during adapter training.
Explore evaluation patternsSee how to apply these findings
Learn More

Key Data

1Adapter-based transfer updates only 3.5–7.3% of the parameters compared to full-policy multi-agent training.
2Evaluated on 3 tasks (pathfinding, navigation, cooperative discovery) spanning discrete and continuous settings and consistently outperformed multi-agent training from scratch.
3Adapters trained at one team size retained strong task performance when deployed with different, including substantially larger, agent counts (generalization across unseen team sizes).

Why It Matters

Engineers building multi-agent systems who want to avoid costly retraining can use this to reuse existing single-agent controllers. Technical leads evaluating deployment strategies gain a low-risk option to preserve proven agent behavior while enabling coordination. Researchers can use the adapter idea to study modular transfer between single- and multi-agent settings. This aligns with the Hierarchical Multi-Agent Pattern.

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Limitations

Requires a meaningful single-agent counterpart where the solo observation captures core task competence; if individual skills do not transfer, the adapter has little to adapt. Highly interactive tasks where coordination fundamentally changes the control objective may need more than input-side adaptation. Paper reports results on three benchmark tasks—performance may vary on very different domains or with heterogeneous agents. Care should be taken with guardrails to prevent unintended coordination, per the Guardrails Pattern.

Deep Dive

Train a capable single-agent policy first, then transfer it unchanged into a multi-agent setting by placing a small, trainable observation adapter in front of the frozen controller. The adapter takes the full local observation (including neighbor info) and outputs a transformed observation in the format the solo policy expects. During multi-agent learning only the adapter (and optionally a small value head) is updated; the actor and original critic remain frozen but are kept in the computation graph so gradients flow to the adapter. This mirrors the Planning Pattern. Across three tasks—lifelong pathfinding, navigation, and cooperative discovery—this setup consistently outperformed conventional multi-agent training from scratch and approached the performance of full fine-tuning while updating only 3.5–7.3% of parameters. The method is compatible with on-policy and off-policy algorithms and works with either shared adapters for homogeneous teams or separate adapters per agent. The result is a practical, low-risk way to move existing single-agent policies into team settings, preserving prior capabilities while learning how to account for neighbors with minimal extra cost. It also fits into Market-Based Coordination Pattern concepts.
Not sure where to start?Get personalized recommendations
Learn More
Credibility Assessment:

ArXiv preprint with low h-index authors (h=3) and no known institutional affiliations; signals indicate an emerging/limited profile.