GeoNav: Empowering MLLMs with Explicit Geospatial Reasoning Abilities for Language-Goal Aerial Navigation
2025/04/13 by Xu, Haotian, Hu, Yue, Gao, Chen +4 · 15 citations
#FOS: Computer and information sciences #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2504.09587
Abstract
Language-goal aerial navigation is a critical challenge in embodied AI, requiring UAVs to localize targets in complex environments such as urban blocks based on textual specification. Existing methods, often adapted from indoor navigation, struggle to scale due to limited field of view, semantic ambiguity among objects, and lack of structured spatial reasoning. In this work, we propose GeoNav, a geospatially aware multimodal agent to enable long-range navigation. GeoNav operates in three phases-landmark navigation, target search, and precise localization-mimicking human coarse-to-fine spatial strategies. To support such reasoning, it dynamically builds two different types of spatial memory. The first is a global but schematic cognitive map, which fuses prior textual geographic knowledge and embodied visual cues into a top-down, annotated form for fast navigation to the landmark region. The second is a local but delicate scene graph representing hierarchical spatial relationships between blocks, landmarks, and objects, which is used for definite target localization. On top of this structured representation, GeoNav employs a spatially aware, multimodal chain-of-thought prompting mechanism to enable multimodal large language models with efficient and interpretable decision-making across stages. On the CityNav urban navigation benchmark, GeoNav surpasses the current state-of-the-art by up to 12.53% in success rate and significantly improves navigation efficiency, even in hard-level tasks. Ablation studies highlight the importance of each module, showcasing how geospatial representations and coarse-to-fine reasoning enhance UAV navigation.
Citations
- Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
- Aerial Vision-and-Language Navigation with Grid-based View Selection and Map Construction
- UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
- MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation
- Qwen2.5-VL Technical Report
- CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space
- OpenBench: A New Benchmark and Baseline for Semantic Navigation in Smart Logistics
- Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation
- NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation
- UEVAVD: A Dataset for Developing UAV's Eye View Active Object Detection
- GPT-4o System Card
- Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and Methodology
- NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UAV Search Missions
- UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
- AeroVerse: UAV-Agent Benchmark Suite for Simulating, Pre-training, Finetuning, and Evaluating Aerospace Embodied World Models
- CityNav: A Large-Scale Dataset for Real-World Aerial Navigation
- SpatialBot: Precise Spatial Understanding with Vision Language Models
- GOMAA-Geo: GOal Modality Agnostic Active Geo-localization
- Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models
- CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud Data
- AerialVLN: Vision-and-Language Navigation for UAVs
- VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models
- Micrograph segmentations for DDEVD
- Segment Anything
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- Path-Planning for Unmanned Aerial Vehicles with Environment Complexity Considerations: A Survey
- LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action
- The Parallelism Tradeoff: Limitations of Log-Precision Transformers
- Aerial Vision-and-Dialog Navigation
- SensatUrban: Learning Semantics from Urban-Scale Photogrammetric Point Clouds
- SensatUrban: Learning Semantics from Urban-Scale Photogrammetric Point Clouds
- Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments
- Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation
- Curiosity-driven Exploration by Self-supervised Prediction
- Openfly: A comprehensive platform for aerial vision-language navigation
Cited by
Related