The Big Picture
A simple three-axis framework (task, capability, scale) reveals where spatially-aware agents work well, shows 68% of research focuses on room-scale problems, and highlights six big research gaps that matter for real deployments.
ON THIS PAGE
The Evidence
A unified taxonomy maps what agents do (navigation, scene understanding, manipulation, geospatial analysis) to how they reason and act (memory, planning, tool use) and to the spatial scale (centimeter, room, city). Analysis of 742 cited works (from over 2,000 reviewed) shows a heavy bias toward room-scale navigation and scene tasks, while fine-grained manipulation and city-scale reasoning are underexplored. Key architectural patterns that bridge perception and action are integrating graph-structured spatial memory (graphs) with language models, predictive world models that simulate outcomes, and multimodal perception-to-action pipelines. The survey ends with six grand challenges—cross-scale representations, long-horizon grounded planning, safety guarantees, sim-to-real transfer, multi-agent coordination, and running agents on edge hardware—that define high-impact directions. world models
Data Highlights
168% of methods target meso-spatial tasks (room- or local-scale), revealing a strong research concentration at that scale.
2742 works are directly cited after reviewing over 2,000 papers, indicating broad literature coverage but selective emphasis.
3A vision-language model (GPT-4V) fails on about 40% of spatial relationship questions in a spatial benchmark, highlighting serious spatial reasoning gaps.
What This Means
Engineers building robots or embodied AI should use the taxonomy to pick the right internal representations (e.g., spatial maps vs episodic logs) and to avoid scale mismatches when transferring policies. Technical leaders and product managers can use the identified design patterns and the roadmap to prioritize investments (e.g., world models, graph-based memory, human-in-the-loop validation) where industry needs outstrip research coverage. Researchers can target underexplored regions (micro manipulation and macro geospatial action) that the survey flags as high-value opportunities.
Not sure where to start?Get personalized recommendations
Key Figures

Fig 1: Figure 1: A unified three-axis taxonomy connecting Agentic AI capabilities with Spatial Intelligence domains across spatial scales. The intersection of these dimensions defines the design space for autonomous spatial intelligence systems. Design trade-off : No existing method achieves strong performance across all three axes simultaneously. Micro-scale manipulation systems (bottom-left) achieve centimeter precision but cannot plan beyond the immediate workspace. Geospatial models (top-right) reason at planetary scale but lack closed-loop action capabilities. The sparsely populated regions of this space (e.g., macro-scale manipulation, micro-scale geospatial) represent open research opportunities where cross-axis integration could yield high-impact advances.
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreConsiderations
The work is a survey synthesizing prior results rather than introducing new experimental evidence; conclusions depend on the selected corpus and filtering choices (years 2018–2026, top venues). The 68% figure may partly reflect benchmark availability and community tooling that favors room-scale problems, not only intrinsic importance. Practical systems still face engineering gaps (robust sim-to-real transfer, runtime constraints on edge devices) that are outside the survey’s empirical scope.
Methodology & More
A three-axis taxonomy connects what spatial agents do (navigation, scene understanding, manipulation, geospatial analysis) with how they act (memory, planning, tool use) and the spatial scale (micro: centimeter, meso: room-level, macro: city/planet). The literature search spanned major databases and venues, filtered for 2018–2026, and used a multi-stage relevance and quality process; the final curated set cites 742 works drawn from over 2,000 reviewed papers. Every method in the survey is mapped to all three axes to enable direct cross-method comparisons and to spot underexplored regions of the design space. The analysis shows strong convergence on several architectural patterns that help turn perception into action: graph-structured spatial memory combined with language reasoning to maintain persistent, relational scene state; world models that predict future states for planning; and multimodal perception-to-action pipelines that ground language and vision in control primitives. A pronounced imbalance appears: most research targets meso-scale navigation and scene tasks, leaving micro-scale manipulation and macro-scale geospatial action insufficiently covered. The survey synthesizes these findings into six grand challenges—unified cross-scale representations, grounded long-horizon planning, safety and verification, reliable sim-to-real transfer, multi-agent coordination, and efficient edge deployment—and offers design patterns (including human-in-the-loop validation) to guide both research and industry adoption.
Avoid common pitfallsLearn what failures to watch for
Credibility Assessment:
Authors have no h-index signals or affiliations provided and zero citations — minimal identifiable credibility.