M3DMap: Object-aware Multimodal 3D Mapping for Dynamic Environments
2025/08/23 by Yudin, Dmitry
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2508.17044
Abstract
3D mapping in dynamic environments poses a challenge for modern researchers in robotics and autonomous transportation. There are no universal representations for dynamic 3D scenes that incorporate multimodal data such as images, point clouds, and text. This article takes a step toward solving this problem. It proposes a taxonomy of methods for constructing multimodal 3D maps, classifying contemporary approaches based on scene types and representations, learning methods, and practical applications. Using this taxonomy, a brief structured analysis of recent methods is provided. The article also describes an original modular method called M3DMap, designed for object-aware construction of multimodal 3D maps for both static and dynamic scenes. It consists of several interconnected components: a neural multimodal object segmentation and tracking module; an odometry estimation module, including trainable algorithms; a module for 3D map construction and updating with various implementations depending on the desired scene representation; and a multimodal data retrieval module. The article highlights original implementations of these modules and their advantages in solving various practical tasks, from 3D object grounding to mobile manipulation. Additionally, it presents theoretical propositions demonstrating the positive effect of using multimodal data and modern foundational models in 3D mapping methods. Details of the taxonomy and method implementation are available at https://yuddim.github.io/M3DMap.
Citations
- SGN-CIRL: Scene Graph-based Navigation with Curriculum, Imitation, and Reinforcement Learning
- LEG-SLAM: Real-Time Language-Enhanced Gaussian Splatting for SLAM
- Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
- Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
- Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces
- Sonata: Self-Supervised Learning of Reliable Point Representations
- Universal Scene Graph Generation
- YOLOE: Real-Time Seeing Anything
- SplatTalk: 3D VQA with Gaussian Splatting
- DriveTransformer: Unified Transformer for Scalable End-to-End Autonomous Driving
- GaussianGraph: 3D Gaussian-based Scene Graph Generation for Open-world Scene Understanding
- BEVDriver: Leveraging BEV Maps in LLMs for Robust Closed-Loop Driving
- MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments
- DynamicGSG: Dynamic 3D Gaussian Scene Graphs for Environment Adaptation
- Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
- GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
- 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding
- GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering
- Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
- Sparse Voxels Rasterization: Real-time High-fidelity Radiance Field Rendering
- 3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
- Multiview Scene Graph
- WildFusion: Multimodal Implicit 3D Reconstructions in the Wild
- ReMEmbR: Building and Reasoning Over Long-Horizon Spatio-Temporal Memory for Robot Navigation
- Point2Graph: An End-to-end Point Cloud-based 3D Open-Vocabulary Scene Graph for Robot Navigation
- FAST-LIVO2: Fast, Direct LiDAR-Inertial-Visual Odometry
- BEVPlace++: Fast, Robust, and Lightweight LiDAR Global Localization for Unmanned Ground Vehicles
- 3D Question Answering for City Scene Understanding
- MSSPlace: Multi-Sensor Place Recognition with Visual and Text Semantics
- AriGraph: Learning Knowledge Graph World Models with Episodic Memory for LLM Agents
- Beyond Bare Queries: Open-Vocabulary Object Grounding with 3D Scene Graph
- OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding
- SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
- Grounded 3D-LLM with Referent Tokens
- 4D Panoptic Scene Graph Generation
- OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
- Weakly-Supervised 3D Scene Graph Generation via Visual-Linguistic Assisted Pseudo-labeling
- OFMPNet: Deep End-to-End Model for Occupancy and Flow Prediction in Urban Environment
- SUGAR: Pre-training 3D Visual Representations for Robotics
- Semantic Gaussians: Open-Vocabulary Scene Understanding with 3D Gaussian Splatting
- Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
- Lifelong LERF: Local 3D Semantic Inventory Monitoring Using FogROS2
- OpenGraph: Open-Vocabulary Hierarchical 3D Graph Representation in Large-Scale Outdoor Environments
- SeCG: Semantic-Enhanced 3D Visual Grounding via Cross-modal Graph Attention
- Learning Generalizable Feature Fields for Mobile Manipulation
- Khronos: A Unified Approach for Spatio-Temporal Metric-Semantic SLAM in Dynamic Environments
- Open3DSG: Open-Vocabulary 3D Scene Graphs from Point Clouds with Queryable Objects and Open-Set Relationships
- YOLO-World: Real-Time Open-Vocabulary Object Detection
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- LangSplat: 3D Language Gaussian Splatting
- EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI
- Indoor and Outdoor 3D Scene Graph Generation via Language-Enabled Spatial Ontologies
- Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
- uSF: Learning Neural Semantic Field with Uncertainty
- LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning
- OVIR-3D: Open-Vocabulary 3D Instance Retrieval Without Training on 3D Data
- Neural Potential Field for Obstacle-Aware Local Motion Planning
- Lang3DSG: Language-based contrastive pre-training for 3D Scene Graph prediction
- Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting
- 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
- Uni3D: Exploring Unified 3D Representation at Scale
- Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene Representation
- ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
- SGRec3D: Self-Supervised 3D Scene Graph Learning via Object-Level Scene Reconstruction
- <b>D</b>ataset for <b>O</b>pen <b>V</b>ocabulary <b>E</b>ntity <b>G</b>rounding (DOVE-G)
- Unsupervised 3D Perception with 2D Vision-Language Distillation for Autonomous Driving
- LLM-Grounder: Open-Vocabulary 3D Visual Grounding with Large Language Model as an Agent
- EigenPlaces: Training Viewpoint Robust Models for Visual Place Recognition
- Bird's-Eye-View Scene Graph for Vision-Language Navigation
- AnyLoc: Towards Universal Visual Place Recognition
- SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning
- Modeling Dynamic Environments with Scene Graph Memory
- Weakly Supervised 3D Open-vocabulary Segmentation
- Foundations of Spatial Perception for Robotics: Hierarchical Representations and Real-time Systems
- Incremental 3D Semantic Scene Graph Prediction from RGB Sequences
- Micrograph segmentations for DDEVD
- Segment Anything
- RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding
- Semantic Ray: Learning a Generalizable Semantic Field with Cross-Reprojection Attention
- LABRAD-OR: Lightweight Memory Scene Graphs for Accurate Bimodal Reasoning in Dynamic Operating Rooms
- VAD: Vectorized Scene Representation for Efficient Autonomous Driving
- SGFormer: Semantic Graph Transformer for Point Cloud-based 3D Scene Graph Generation
- LERF: Language Embedded Radiance Fields
- Audio Visual Language Maps for Robot Navigation
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- CLIP-FO3D: Learning Free Open-world 3D Scene Representations from 2D Dense CLIP
- ConceptFusion: Open-set Multimodal 3D Mapping
- Rethinking Voxelization and Classification for 3D Object Detection
- MixVPR: Feature Mixing for Visual Place Recognition
- HPointLoc: Point-based Indoor Place Recognition using Synthetic RGB-D Images
- Planning-oriented Autonomous Driving
- PLA: Language-Driven Open-Vocabulary 3D Scene Understanding
- OpenScene: 3D Scene Understanding with Open Vocabularies
- Language Conditioned Spatial Relation Reasoning for 3D Object Grounding
- Multi-Object Navigation with dynamically learned neural implicit representations
- Visual Language Maps for Robot Navigation
- Open-vocabulary Queryable Scene Representations for Real World Planning
- BEVerse: Unified Perception and Prediction in Birds-Eye-View for Vision-Centric Autonomous Driving
- Rethinking Visual Geo-localization for Large-Scale Applications
- Improving Point Cloud Based Place Recognition with Ranking-based Loss and Large Batch Training
- Language-driven Semantic Segmentation
- AdaFusion: Visual-LiDAR Fusion with Adaptive Weights for Place Recognition
- Hierarchical Representations and Explicit Memory: Learning Effective Navigation Policies on 3D Scene Graphs using Graph Neural Networks
- TransLoc3D : Point Cloud based Large-scale Place Recognition using Adaptive Receptive Fields
- SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition
- Hierarchical Cross-Modal Agent for Robotics Vision-and-Language Navigation
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion Dataset
- MinkLoc++: Lidar and Monocular Image Fusion for Place Recognition
- In-Place Scene Labelling and Understanding with Implicit Scene Representation
- SceneGraphFusion: Incremental 3D Scene Graph Prediction from RGB-D Sequences
- NDT-Transformer: Large-Scale 3D Point Cloud Localisation using the Normal Distribution Transform Representation
- Exploiting Edge-Oriented Reasoning for 3D Point-based Scene Graph Analysis
- Learning Transferable Visual Models From Natural Language Supervision
- 3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and Camera
- RIO: 3D Object Instance Re-Localization in Changing Indoor Environments
- PointNetVLAD: Deep Point Cloud Based Retrieval for Large-Scale Place\n Recognition
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- NetVLAD: CNN architecture for weakly supervised place recognition
- Gemini Robotics: Bringing AI into the Physical World
- Mapping the Unseen: Unified Promptable Panoptic Mapping with Dynamic Labeling using Foundation Models
- 3D-LLM: Injecting the 3D World into Large Language Models
Related