Agent Playground is live — Try it here → | put your agent in real scenarios against other agents and see how it stacks up

The Big Picture

A configuration-based planner adapted from multi-robot pathfinding lets teams of agents coordinate pushes and solve multi-box moving tasks quickly, reliably, and often optimally—scaling beyond single-agent search methods.

The Evidence

Adapting a configuration-space search backbone plus a coordinated action generator produces a solver (Sokoban-LaCAM) that finds feasible multi-agent box-moving plans almost instantly and often proves optimal within practical time limits. The approach handles agent–agent and box collisions, integrates simple task assignment heuristics, and is complete (it will find a solution if one exists) and eventually optimal under cumulative cost. Compared with a baseline that extends single-agent Sokoban search, the configuration-based method solves more instances and scales much better. planning pattern
Not sure where to start?Get personalized recommendations
Learn More

Data Highlights

1Evaluation used 90 randomized instances on a 5×5 map (10 instances per setting) with up to 6 agents and 6 boxes and a 30 s planning limit.
2Sokoban-LaCAM produced initial feasible plans nearly instantly and converged to optimal solutions on every instance that the adapted IDA* baseline could solve, often proving optimality within the 30 s limit.
3The method is claimed to solve instances with 'tens' of agents and boxes within seconds, demonstrating scalability beyond traditional single-agent Sokoban solvers.

What This Means

Engineers building coordinated robot fleets (warehouse automation, pallet movers) can use a configuration-based planning backbone to combine task assignment and collision-free movement. Technical leads and researchers evaluating multi-agent orchestration should note this as a practical, extendable primitive that scales better than single-agent search baselines and can be augmented with learned heuristics. planning pattern

Key Figures

Figure 1 : Planning performance of Sokoban-LaCAM and its underlying configuration generator, Sokoban-PIBT. The top row illustrates representative problem instances. The bottom rows report the success rate under a 10 s 10\text{\,}\mathrm{s} time limit across different maps and varying numbers of agents and boxes. Each scenario consists of ten instances. The centre-rightmost plot shows the aggregated success rate as a function of elapsed planning time.
Fig 1: Figure 1 : Planning performance of Sokoban-LaCAM and its underlying configuration generator, Sokoban-PIBT. The top row illustrates representative problem instances. The bottom rows report the success rate under a 10 s 10\text{\,}\mathrm{s} time limit across different maps and varying numbers of agents and boxes. Each scenario consists of ten instances. The centre-rightmost plot shows the aggregated success rate as a function of elapsed planning time.

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Yes, But...

Experiments are primarily on randomly generated grid maps and small standard maps, so performance on realistic, structured warehouse layouts may differ. Some engineering details and parameter choices are omitted, so reproducing top performance may require tuning. The Sokoban rules (push-only box moves, equal numbers of boxes and targets) simplify some real-world constraints; adapting to richer manipulation actions will need additional work. emergence-aware monitoring pattern

Methodology & More

Sokoban-LaCAM brings a configuration-space multi-agent search backbone to the puzzle of coordinating many agents that push boxes to targets. At each step the planner searches over full system configurations (positions of agents and boxes) while generating joint agent actions using a modified preference-based local policy that respects pushing rules and rejects actions leading to box–box or box–agent collisions. Because boxes only move when pushed, the generator focuses on agent actions; the overall search retains theoretical guarantees of completeness and eventual optimality for cumulative-cost objectives such as plan length. In experiments the solver outperformed an adapted single-agent search baseline (IDA* with Sokoban heuristics): it finds initial feasible plans almost immediately, solves more instances within a 30 s limit, and proves optimality for the instances the baseline could solve. The work shows that established multi-agent pathfinding techniques can serve as a general, scalable primitive for problems that combine task assignment, object manipulation, and coordinated motion. A clear next step is adding learned heuristics or task-specific configuration generators to further speed search and handle richer, real-world manipulation constraints. Blackboard Pattern Sub-Agent Delegation Pattern
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

Single author with very low h‑index and no listed affiliation or venue beyond arXiv — lacks identifiable institutional or reputational signals.