Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
2025/11/17 by Lingfeng Zhang, Yuchen Zhang, Zhang, Lingfeng +16
Computer Science · Engineering · #Multimodal Machine Learning Applications #Robotics and Sensor-Based Localization #Advanced Neural Network Applications
paper · pdf · doi:10.48550/arxiv.2511.13269
Abstract
Vision-Language Models (VLMs), leveraging their powerful visual perception and reasoning capabilities, have been widely applied in Unmanned Aerial Vehicle (UAV) tasks. However, the spatial intelligence capabilities of existing VLMs in UAV scenarios remain largely unexplored, raising concerns about their effectiveness in navigating and interpreting dynamic environments. To bridge this gap, we introduce SpatialSky-Bench, a comprehensive benchmark specifically designed to evaluate the spatial intelligence capabilities of VLMs in UAV navigation. Our benchmark comprises two categories-Environmental Perception and Scene Understanding-divided into 13 subcategories, including bounding boxes, color, distance, height, and landing safety analysis, among others. Extensive evaluations of various mainstream open-source and closed-source VLMs reveal unsatisfactory performance in complex UAV navigation scenarios, highlighting significant gaps in their spatial capabilities. To address this challenge, we developed the SpatialSky-Dataset, a comprehensive dataset containing 1M samples with diverse annotations across various scenarios. Leveraging this dataset, we introduce Sky-VLM, a specialized VLM designed for UAV spatial reasoning across multiple granularities and contexts. Extensive experimental results demonstrate that Sky-VLM achieves state-of-the-art performance across all benchmark tasks, paving the way for the development of VLMs suitable for UAV scenarios. The source code is available at https://github.com/linglingxiansen/SpatialSKy.
Citations
- MiMo-Embodied: X-Embodied Foundation Model Technical Report
- SoraNav: Adaptive UAV Task-Centric Navigation via Zeroshot VLM Reasoning
- Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4
- See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation
- VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results
- TopoNav: Topological Graphs as a Key Enabler for Advanced Object Navigation
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- RynnEC: Bringing MLLMs into Embodied World
- NavA3: Understanding Any Instruction, Navigating Anywhere, Finding Anything
- UAVScenes: A Multi-Modal Dataset for UAVs
- SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban Environments
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- RoboBrain 2.0 Technical Report
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
- Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation
- VLM-RRT: Vision Language Model Guided RRT Search for Autonomous UAV Navigation
- Qwen3 Technical Report
- SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
- Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
- Exploring Typographic Visual Prompts Injection Threats in Cross-Modality Generation Models
- HumanoidPano: Hybrid Spherical Panoramic-LiDAR Cross-Modal Perception for Humanoid Robots
- LiPS: Large-Scale Humanoid Robot Reinforcement Learning with Parallel-Series Structures
- UAV-VLRR: Vision-Language Informed NMPC for Rapid Response in UAV Search and Rescue
- AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter
- RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
- Qwen2.5-VL Technical Report
- MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation
- Embodied Scene Understanding for Vision Language Models via MetaVQA
- UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
- GPT-4o System Card
- Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and Methodology
- Multi-Floor Zero-Shot Object Navigation Policy
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- TriHelper: Zero-Shot Object Navigation with Dynamic Assistance
- The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- AerialVLN: Vision-and-Language Navigation for UAVs
- CAD2RL: Real Single-Image Flight without a Single Real Image
- VQA: Visual Question Answering
- Learning to Navigate Socially Through Proactive Risk Perception
- Stairway to Success: An Online Floor-Aware Zero-Shot Object-Goal Navigation Framework via LLM-Driven Coarse-to-Fine Exploration
- Gemini Robotics: Bringing AI into the Physical World
- MindCube: Spatial Mental Modeling from Limited Views
Related