OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models
2025/09/30 by Biswas, Subrata, Khan, Mohammad Nur Hossain, Islam, Bashima · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Sound (cs.SD)
paper · doi:10.48550/arxiv.2509.26140
Abstract
Spatial reasoning is fundamental to auditory perception, yet current audio large language models (ALLMs) largely rely on unstructured binaural cues and single step inference. This limits both perceptual accuracy in direction and distance estimation and the capacity for interpretable reasoning. Recent work such as BAT demonstrates spatial QA with binaural audio, but its reliance on coarse categorical labels (left, right, up, down) and the absence of explicit geometric supervision constrain resolution and robustness. We introduce the Spatial-Acoustic Geometry Encoder (SAGE), a geometry-aware audio encoder that aligns binaural acoustic features with 3D spatial structure using panoramic depth images and room-impulse responses at training time, while requiring only audio at inference. Building on this representation, we present OWL, an ALLM that integrates SAGE with a spatially grounded chain-of-thought to rationalize over direction-of-arrivals (DoA) and distance estimates. Through curriculum learning from perceptual QA to multi-step reasoning, OWL supports o'clock-level azimuth and DoA estimation. To enable large-scale training and evaluation, we construct and release BiDepth, a dataset of over one million QA pairs combining binaural audio with panoramic depth images and room impulse responses across both in-room and out-of-room scenarios. Across two benchmark datasets, our new BiDepth and the public SpatialSoundQA, OWL reduces mean DoA error by \textbf11∘ through SAGE and improves spatial reasoning QA accuracy by up to 25% over BAT.
Citations
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language
- Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
- LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
- Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
- LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
- Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
- Distill Visual Chart Reasoning Ability from LLMs to MLLMs
- LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
- GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
- LLMSense: Harnessing LLMs for High-level Reasoning Over Spatiotemporal Sensor Traces
- Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
- BAT: Learning to Reason about Spatial Sounds with Large Language Models
- Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
- IMUGPT 2.0: Language-Based Cross Modality Transfer for Sensor-Based Human Activity Recognition
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- Improved Baselines with Visual Instruction Tuning
- Qwen Technical Report
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Pengi: An Audio Language Model for Audio Tasks
- Listen, Think, and Understand
- AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
- Visual Instruction Tuning
- GPT-4 Technical Report
- LOCUS: LOcalization with Channel Uncertainty and Sporadic Energy
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
- Masked Autoencoders that Listen
- SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning
- ACCDOA: Activity-Coupled Cartesian Direction of Arrival Representation for Sound Event Localization and Detection
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Matterport3D: Learning from RGB-D Data in Indoor Environments
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary\n Visual Reasoning
- SGDR: Stochastic Gradient Descent with Warm Restarts
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
Cited by
Related