At a Glance
Let the best-performing agent share its strategy and rebuild who listens to whom for each problem — teams of AI agents reason better without any extra training or external judge.
ON THIS PAGE
What They Found
Selecting a single agent as a temporary “strategy donor” based on early answers, copying its role instructions into other agents (while keeping their specialties), and dynamically rebuilding a sparse communication graph yields consistent gains on reasoning and knowledge tasks. The selection relies on simple signals (answer agreement, prefix consistency, and reciprocal peer review) and the communication graph is updated after every collaboration round. The system is training-free and works across different model backbones and both text and vision-language inputs. Dynamic Task Routing Pattern.
Not sure where to start?Get personalized recommendations
Data Highlights
1Evaluated across 6 benchmarks: GSM8K, MATH, GSM-Hard, AQuA-RAT, MMLU, and GPQA-Diamond.
2Main experiments used two backbones with roughly 2.5 billion and 3 billion parameter models (Qwen2.5-1.5B-Instruct and Mistral-3-3B-Instruct).
3Some agent roles were picked as strategy donors more often than chance (chance = 25%); role preference was statistically significant (permutation test p < 0.05) on marked benchmarks.
What This Means
Engineers building multi-agent workflows who want better reliability without retraining models will find this useful. Technical leaders evaluating agent orchestration or agent trust can use the donor-and-graph approach to improve collaboration and spot which agents drive reasoning quality. See the Agent Registry Pattern for related orchestration patterns.
Key Figures

Fig 5: Figure 6: Donor rate by role. Percentage of problems in which a role is selected as donor when it is on the team (chance = 25 % =25\% ), pooled over three runs. Problems decided by the agent-index tie-break are excluded. ∗ marks benchmarks with a significant role preference (permutation test, p < 0.05 p<0.05 ); the other columns are faded.
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreYes, But...
Donor selection signals (agreement, prefix match, peer review) are imperfect and can reinforce shared mistakes if several agents share the same blind spot. Results are reported for fixed teams and the tested model sizes; benefits may change with much larger or very different models. Applications that affect real-world decisions still need independent validation and guardrails since coordinated agents can amplify biased or incorrect reasoning. This aligns with concerns highlighted in Inter-Agent Miscommunication.
Methodology & More
SAGE (Self-Adapting Group of Experts) turns a group of AI agents into a problem-specific, adaptive team without training any models. For each problem, agents first answer independently; then SAGE picks a strategy donor using three lightweight signals: whether answers agree, whether the donor’s answer prefix looks consistent, and how peers rate each other through reciprocal review. The donor’s original role instructions (the way it is asked to reason) are then used to rewrite other agents’ prompts while preserving their specializations. Finally, agents exchange messages along a sparse directed graph that is rebuilt after every collaboration round, so stronger agents influence weaker ones more directly. ReAct Pattern (Reason + Act) and related safety considerations can further inform how prompts are shaped, and the approach resonates with guardrail-focused design such as the Red Teaming Pattern. Experiments ran on six established benchmarks spanning math, knowledge, and scientific reasoning and used two mid-size language models as backbones. Compared to a single-call baseline and several multi-agent methods, the approach consistently improved team reasoning across text and vision-language inputs. The biggest practical benefit is that teams adapt at inference time: a role that happens to reason most sensibly for a given problem becomes a temporary mentor, and communication patterns evolve so useful reasoning flows where it’s needed. That makes SAGE a straightforward add-on for teams that want better, problem-aware collaboration without collecting new data or retraining models.
Avoid common pitfallsLearn what failures to watch for
Credibility Assessment:
Authors include researchers with h-indices in the 10–20 range (e.g., h≈14 and h≈10), indicating established researchers, but no top venue or high-profile affiliations and it's an arXiv preprint — moderate credibility.