At a Glance
Match the team structure to the task: grouping agents for parallel work and stacking managers for ordered steps makes heterogeneous robot teams complete complex missions faster, with better outcomes and competitive compute use.
ON THIS PAGE
Core Insights
Task-specific organizational hierarchies that mirror how work is interdependent (parallel vs. sequential) let many different agents cooperate far more effectively than one-size-fits-all arrangements. A framework that creates horizontal managers for concurrent tasks and vertical managers for ordered phases produced faster progress, higher final scores, and lower communication overhead across diverse wildfire response missions. Automatically generated hierarchies from language models were helpful, but human-designed hierarchies remained the strongest; still, both variants beat four strong prior multi-agent baselines. Orchestrator-Worker Pattern Explainability
By the Numbers
1Evaluated on 25 wildfire-response missions with 8 different language models and 5 random seeds — over 1,000 runs — including scenarios with up to 50 agents.
2ORCH variants outperformed all four baseline coordination methods on final task outcome, execution efficiency (area-under-progress curve), and output-token efficiency across the benchmark.
3ORCH ranked among the top three methods for exploration, input-token usage, and API-call usage while also recording the largest share of first-place finishes across the LLM-task-seed combinations.
Implications
Engineers building multi-agent, embodied systems should care because organizing agents by how work is interdependent can reduce wasted coordination and improve mission success without extra model cost. Technical leaders evaluating agent fleets or choosing language model sizes should note that better organization can let smaller or mid-sized models outperform larger ones in collective settings. Multi-Agent IT Operations
Test your agentsValidate against real scenarios
Key Figures

Fig 1: Figure 1 : Method Overall . (A) The team hierarchy was designed based on the task descriptions and available worker agents, either by a human expert or an LLM with a critic agent. (B) During execution, the manager generates the plan for its worker agents; the worker agents execute the plan and provide feedback for the manager to update the plan. (C) Pooled interdependency refers to a group of agents working concurrently and independently to achieve a goal. For instance, 3 firefighters each cut specific target trees (yellow color) to achieve the goal of cutting a group of target trees. (D) Sequential interdependency refers to a group of agents working sequentially to achieve a goal. For instance, to extinguish a fire, the firefighters must first scout the location of the fire; then the helicopter picks up and drops off the firefighters near the fire; finally, the firefighters extinguish the fire. (E) ORCH enforces hybrid collaboration mode (pooled and sequential interdependency).

Fig 2: Figure 2 : Normalized Aggregated Performance on 25 CREW-Wildfire Tasks, Ranking Analysis, and ANOVA Test Results . The detailed normalization process can be found in Section S3 – S4 in Supplementary Text. (A) ORCH with human-expert-designed and LLM-generated team hierarchies both significantly outperform the other four baselines in terms of output token usage, final score, and efficiency, and they rank top-3 in terms of exploration, input token usage, and the number of API calls. (B) ORCH has the highest proportion of first-place rankings across all the LLM-task-seed combinations for both final score and efficiency metric. (C) The ANOVA test results showed that the interaction term between algorithm and LLM is not statistically significant, while the remaining terms are.

Fig 3: Figure 3 : Normalized Aggregated Performance on 25 CREW-Wildfire Tasks by Language Models . The detailed normalization process can be found in Section S3 – S4 in the Supplementary Text. (A) Aggregated LLM performance on the 25 CREW-Wildfire tasks with ORCH with human-expert-designed team hierarchies. (B) Aggregated LLM performance on the 25 CREW-Wildfire tasks with ORCH with LLM-generated team hierarchies. Surprisingly, the middle-sized models perform better than large-sized models with ORCH under both organizational designs.

Fig 4: Figure 4 : Ablation study on Organization Design, Team Hierarchy Analysis, and Failure Reasons Analysis through Chat History . (A) ORCH consistently outperformed prior baselines across human-designed and LLM-generated team hierarchies, including those generated without critic supervision. (B) We analyzed the team hierarchies generated by different methods. The results show that human-designed hierarchies exhibit the most balanced team structures, followed by LLM-generated hierarchies with critic supervision, while hierarchies generated without critic supervision are less balanced. (C) We classified the failure modes of different algorithms into five categories. For the four prior methods, failures were predominantly attributed to worker-agent execution. In contrast, failures in ORCH were distributed across worker execution, incorrect task allocation, and insufficient replanning, indicating a reduced concentration of failures at the worker-execution level with the principle of organizational design.
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreLimitations
Results come from a 25-task wildfire benchmark; transfer to other domains (logistics, construction, urban search) is plausible but not proven. ORCH builds the hierarchy before execution and adapts plans within that structure — it does not yet restructure the hierarchy dynamically during a mission. Human-designed hierarchies still outperformed automatic generation, so fully automated organizational design needs more work and validation of failure labels produced by language models is required for high-stakes use. Supervisor Pattern
Deep Dive
ORCH (Organizing Roles and Coordination Hierarchies) turns ideas from human organizations into practical coordination rules for embodied agent teams. It classifies interdependence into two actionable types: pooled (independent, parallel work) handled by horizontal managers that allocate concurrent tasks, and sequential (ordered dependencies) handled by vertical managers that enforce phases and handoffs. Given a mission description and a roster of heterogeneous worker agents, ORCH generates a rooted team hierarchy (either designed by an expert or proposed by a language model with critic feedback) and executes mission plans via managers that produce and refine sub-plans for their workers. Reflection Pattern Emergence-Aware Monitoring Pattern Reflection Pattern Byzantine-Resilient Consensus Pattern Agent Registry Pattern Mutual Verification Pattern Emergence-Aware Monitoring Pattern Reflection Pattern [Mirror]
Need expert guidance?We can help implement this
Credibility Assessment:
ArXiv preprint, no listed affiliations or citations and authors not widely known — minimal reputation signals.