ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual–Inertial, and Multimap SLAM
2021/05/25 by Carlos Campos, Richard Elvira, Juan J. Gomez Rodriguez +2 · 101 citations
Engineering · Computer Science · Earth and Planetary Sciences · #Robotics and Sensor-Based Localization #Advanced Vision and Imaging #3D Surveying and Cultural Heritage
paper · pdf · doi:10.1109/tro.2021.3075644
openalex created_date 2020/07/29 · openalex publication_date 2021/05/25 · openalex updated_date 2026/07/31
Abstract
This article presents ORB-SLAM3, the first system able to perform visual, visual-inertial and multimap SLAM with monocular, stereo and RGB-D cameras, using pin-hole and fisheye lens models. The first main novelty is a tightly integrated visual-inertial SLAM system that fully relies on maximuma posteriori(MAP) estimation, even during IMU initialization, resulting in real-time robust operation in small and large, indoor and outdoor environments, being two to ten times more accurate than previous approaches. The second main novelty is a multiple map system relying on a new place recognition method with improved recall that lets ORB-SLAM3 survive to long periods of poor visual information: when it gets lost, it starts a new map that will be seamlessly merged with previous maps when revisiting them. Compared with visual odometry systems that only use information from the last few seconds, ORB-SLAM3 is the first system able to reuse in all the algorithm stages all previous information from high parallax co-visible keyframes, even if they are widely separated in time or come from previous mapping sessions, boosting accuracy. Our experiments show that, in all sensor configurations, ORB-SLAM3 is as robust as the best systems available in the literature and significantly more accurate. Notably, our stereo-inertial SLAM achieves an average accuracy of 3.5 cm in the EuRoC drone and 9 mm under quick hand-held motions in the room of TUM-VI dataset, representative of AR/VR scenarios. For the benefit of the community we make public the source code.
Cited by
- Understanding and Optimizing Attention-Based Sparse Matching for Diverse Local Features
- GeVI-SLAM: Gravity-Enhanced Stereo Visua Inertial SLAM for Underwater Robots
- GSOMAR: a geospatial SLAM framework for real-time outdoor mobile augmented reality scene modeling and semantic interaction
- Analysis and design framework for the development of indoor scene understanding assistive solutions for the person with visual impairment/blindness
- Underwater Visual-Inertial-Acoustic-Depth SLAM with DVL Preintegration for Degraded Environments
- Freehand 3D Ultrasound Imaging: Sim-in-the-Loop Probe Pose Optimization via Visual Servoing
- Deep Learning-Powered Visual SLAM Aimed at Assisting Visually Impaired Navigation
- SEA: Semantic Map Prediction for Active Exploration of Uncertain Areas
- MRASfM: Multi-Camera Reconstruction and Aggregation through Structure-from-Motion in Driving Scenes
- VAR-SLAM: Visual Adaptive and Robust SLAM for Dynamic Environments
- SPORTS: Simultaneous Panoptic Odometry, Rendering, Tracking and Segmentation for Urban Scenes Understanding
- E-MoFlow: Learning Egomotion and Optical Flow from Event Data via Implicit Regularization
- Visual Odometry with Transformers
- Fast Vision in the Dark: A Case for Single-Photon Imaging in Planetary Navigation
- EC3R-SLAM: Efficient and Consistent Monocular Dense SLAM with Feed-Forward 3D Reconstruction
- sqrtVINS: Robust and Ultrafast Square-Root Filter-based 3D Motion Tracking
- VG-Mapping: Variation-Aware 3D Gaussians for Online Semi-static Scene Mapping
- Zero-shot Structure Learning and Planning for Autonomous Robot Navigation using Active Inference
- Robust Visual Teach-and-Repeat Navigation with Flexible Topo-metric Graph Map Representation
- Non-Rigid Structure-from-Motion via Differential Geometry with Recoverable Conformal Scale
- Dropping the D: RGB-D SLAM Without the Depth Sensor
- CLEAR-IR: Clarity-Enhanced Active Reconstruction of Infrared Imagery
- OKVIS2-X: Open Keyframe-based Visual-Inertial SLAM Configurable with Dense Depth or LiDAR, and GNSS
- Statistical Uncertainty Learning for Robust Visual-Inertial State Estimation
- TCB-VIO: Tightly-Coupled Focal-Plane Binary-Enhanced Visual Inertial Odometry
- Kilometer-Scale GNSS-Denied UAV Navigation via Heightmap Gradients: A Winning System from the SPRIN-D Challenge
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- Non-submodular Visual Attention for Robot Navigation
- Benchmarking Egocentric Visual-Inertial SLAM at City Scale
- Graphite: A GPU-Accelerated Mixed-Precision Graph Optimization Framework
- User-Centric Communication Service Provision for Edge-Assisted Mobile Augmented Reality
- Advancing Data Quality of Marine Archaeological Documentation Using Underwater Robotics: From Simulation Environments to Real-World Scenarios
- Good Weights: Proactive, Adaptive Dead Reckoning Fusion for Continuous and Robust Visual SLAM
- Real-Time Indoor Object SLAM with LLM-Enhanced Priors
- PL-VIWO2: A Lightweight, Fast and Robust Visual-Inertial-Wheel Odometry Using Points and Lines
- MASt3R-Fusion: Integrating Feed-Forward Visual Model with IMU, GNSS for High-Functionality SLAM
- SLAM-Free Visual Navigation with Hierarchical Vision-Language Perception and Coarse-to-Fine Semantic Topological Planning
- Deep learning-based 3D morphological segmentation and quantitative growth analysis of field-grown cabbage across the full cycle
- SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization
- Crater Observing Bio-inspired Rolling Articulator (COBRA)
- MV-UMI: A Scalable Multi-View Interface for Cross-Embodiment Learning
- HUNT: High-Speed UAV Navigation and Tracking in Unstructured Environments via Instantaneous Relative Frames
- Combining Projected Uncertainty for Self-Supervised Visual Odometry: From Two-Frame to Multi-Frame
- ProDyG: Progressive Dynamic Scene Reconstruction via Gaussian Splatting from Monocular Videos
- SLAM-Former: Putting SLAM into One Transformer
- See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
- Omni-LIVO: Robust RGB-Colored Multi-Camera Visual-Inertial-LiDAR Odometry via Photometric Migration and ESIKF Fusion
- BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots
- MCGS-SLAM: A Multi-Camera SLAM Framework Using Gaussian Splatting for High-Fidelity Mapping
- BIM Informed Visual SLAM for Construction Monitoring
- FlightDiffusion: Revolutionising Autonomous Drone Training with Diffusion Models Generating FPV Video
- ROOM: A Physics-Based Continuum Robot Simulator for Photorealistic Medical Datasets Generation
- Maps for Autonomous Driving: Full-process Survey and Frontiers
- VROOM - Visual Reconstruction over Onboard Multiview
- MemGS: Memory-Efficient Gaussian Splatting for Real-Time SLAM
- FastTrack: GPU-Accelerated Tracking for Visual SLAM
- Efficient and Accurate Downfacing Visual Inertial Odometry
- SMapper: A Multi-Modal Data Acquisition Platform for SLAM Benchmarking
- Online Dynamic SLAM with Incremental Smoothing and Mapping
- Radar-Based Odometry for Low-Speed Driving
- Real-time Photorealistic Mapping for Situational Awareness in Robot Teleoperation
- Evaluating Magic Leap 2 Tool Tracking for AR Sensor Guidance in Industrial Inspections
- IL-SLAM: Intelligent Line-assisted SLAM Based on Feature Awareness for Dynamic Environments
- Articulated Object Estimation in the Wild
- ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association
- SR-SLAM: Scene-reliability Based RGB-D SLAM in Diverse Environments
- FGO-SLAM: Enhancing Gaussian SLAM with Globally Consistent Opacity Radiance Field
- DyPho-SLAM : Real-time Photorealistic SLAM in Dynamic Environments
- Tank dataset: An underwater multi-sensor dataset for SLAM evaluation
- The Rosario Dataset v2: Multimodal Dataset for Agricultural Robotics
- UnPose: Uncertainty-Guided Diffusion Priors for Zero-Shot Pose Estimation
- HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation
- Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey
- EAROL: Environmental Augmented Perception-Aware Planning and Robust Odometry via Downward-Mounted Tilted LiDAR
- ROVER: Robust Loop Closure Verification with Trajectory Prior in Repetitive Environments
- Unifying Scale-Aware Depth Prediction and Perceptual Priors for Monocular Endoscope Pose Estimation and Tissue Reconstruction
- ViPE: Video Pose Engine for 3D Geometric Perception
- QoE-Aware Service Provision for Mobile AR Rendering: An Agent-Driven Approach
- XR Reality Check: What Commercial Devices Deliver for Spatial Tracking
- False Reality: Uncovering Sensor-induced Human-VR Interaction Vulnerability
- Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction
- Autonomous Navigation of Cloud-Controlled Quadcopters in Confined Spaces Using Multi-Modal Perception and LLM-Driven High Semantic Reasoning
- Bio-Inspired Topological Autonomous Navigation with Active Inference in Robotics
- Navigation and Exploration with Active Inference: from Biology to Industry
- EGS-SLAM: RGB-D Gaussian Splatting SLAM with Events
- RiemanLine: Riemannian Manifold Representation of 3D Lines for Factor Graph Optimization
- Opti-Acoustic Scene Reconstruction in Highly Turbid Underwater Environments
- SGAD: Semantic and Geometric-aware Descriptor for Local Feature Matching
- LT-Gaussian: Long-Term Map Update Using 3D Gaussian Splatting for Autonomous Driving
- CoProU-VO: Combining Projected Uncertainty for End-to-End Unsupervised Monocular Visual Odometry
- The Monado SLAM Dataset for Egocentric Visual-Inertial Tracking
- Stereo 3D Gaussian Splatting SLAM for Outdoor Urban Scenes
- DuLoc: Life-Long Dual-Layer Localization in Changing and Dynamic Expansive Scenarios
- Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
- Adaptive Prior Scene-Object SLAM for Dynamic Environments
- FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
- FROSS: Faster-than-Real-Time Online 3D Semantic Scene Graph Generation from RGB-D Images
- DOA: A Degeneracy Optimization Agent with Adaptive Pose Compensation Capability based on Deep Reinforcement Learning
- G2S-ICP SLAM: Geometry-aware Gaussian Splatting ICP SLAM
- FAST-LIO2: Fast Direct LiDAR-Inertial Odometry
- <scp>3D LiDAR SLAM</scp>: A survey
Related