ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM
2020/07/31 by Carlos Campos, Richard Elvira, Juan J. Gómez Rodríguez +5 · 208 citations
Computer Science · Earth and Planetary Sciences · Engineering · #3D Surveying and Cultural Heritage #Advanced Vision and Imaging #Robotics and Sensor-Based Localization #cs.RO
paper · pdf · doi:10.1109/tro.2021.3075644
openalex created_date 2020/07/29 · arxiv created 2021/04/23 · arxiv updated 2021/04/26 · openalex publication_date 2021/05/25 · openalex updated_date 2026/08/04
Abstract
This paper presents ORB-SLAM3, the first system able to perform visual, visual-inertial and multi-map SLAM with monocular, stereo and RGB-D cameras, using pin-hole and fisheye lens models. The first main novelty is a feature-based tightly-integrated visual-inertial SLAM system that fully relies on Maximum-a-Posteriori (MAP) estimation, even during the IMU initialization phase. The result is a system that operates robustly in real-time, in small and large, indoor and outdoor environments, and is 2 to 5 times more accurate than previous approaches. The second main novelty is a multiple map system that relies on a new place recognition method with improved recall. Thanks to it, ORB-SLAM3 is able to survive to long periods of poor visual information: when it gets lost, it starts a new map that will be seamlessly merged with previous maps when revisiting mapped areas. Compared with visual odometry systems that only use information from the last few seconds, ORB-SLAM3 is the first system able to reuse in all the algorithm stages all previous information. This allows to include in bundle adjustment co-visible keyframes, that provide high parallax observations boosting accuracy, even if they are widely separated in time or if they come from a previous mapping session. Our experiments show that, in all sensor configurations, ORB-SLAM3 is as robust as the best systems available in the literature, and significantly more accurate. Notably, our stereo-inertial SLAM achieves an average accuracy of 3.6 cm on the EuRoC drone and 9 mm under quick hand-held motions in the room of TUM-VI dataset, a setting representative of AR/VR scenarios. For the benefit of the community we make public the source code.
Citations
Cited by
- Understanding and Optimizing Attention-Based Sparse Matching for Diverse Local Features
- GeVI-SLAM: Gravity-Enhanced Stereo Visua Inertial SLAM for Underwater Robots
- GSOMAR: a geospatial SLAM framework for real-time outdoor mobile augmented reality scene modeling and semantic interaction
- Analysis and design framework for the development of indoor scene understanding assistive solutions for the person with visual impairment/blindness
- Underwater Visual-Inertial-Acoustic-Depth SLAM with DVL Preintegration for Degraded Environments
- Freehand 3D Ultrasound Imaging: Sim-in-the-Loop Probe Pose Optimization via Visual Servoing
- Deep Learning-Powered Visual SLAM Aimed at Assisting Visually Impaired Navigation
- SEA: Semantic Map Prediction for Active Exploration of Uncertain Areas
- MRASfM: Multi-Camera Reconstruction and Aggregation through Structure-from-Motion in Driving Scenes
- VAR-SLAM: Visual Adaptive and Robust SLAM for Dynamic Environments
- SPORTS: Simultaneous Panoptic Odometry, Rendering, Tracking and Segmentation for Urban Scenes Understanding
- E-MoFlow: Learning Egomotion and Optical Flow from Event Data via Implicit Regularization
- FVO: Fast Visual Odometry with Transformers
- Fast Vision in the Dark: A Case for Single-Photon Imaging in Planetary Navigation
- EC3R-SLAM: Efficient and Consistent Monocular Dense SLAM with Feed-Forward 3D Reconstruction
- sqrtVINS: Robust and Ultrafast Square-Root Filter-based 3D Motion Tracking
- VG-Mapping: Variation-Aware 3D Gaussians for Online Semi-static Scene Mapping
- Zero-shot Structure Learning and Planning for Autonomous Robot Navigation using Active Inference
- Robust Visual Teach-and-Repeat Navigation with Flexible Topo-metric Graph Map Representation
- Non-Rigid Structure-from-Motion via Differential Geometry with Recoverable Conformal Scale
- Dropping the D: RGB-D SLAM Without the Depth Sensor
- CLEAR-IR: Clarity-Enhanced Active Reconstruction of Infrared Imagery
- OKVIS2-X: Open Keyframe-based Visual-Inertial SLAM Configurable with Dense Depth or LiDAR, and GNSS
- Statistical Uncertainty Learning for Robust Visual-Inertial State Estimation
- TCB-VIO: Tightly-Coupled Focal-Plane Binary-Enhanced Visual Inertial Odometry
- Kilometer-Scale GNSS-Denied UAV Navigation via Heightmap Gradients: A Winning System from the SPRIN-D Challenge
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- Non-submodular Visual Attention for Robot Navigation
- Benchmarking Egocentric Visual-Inertial SLAM at City Scale
- Graphite: A GPU-Accelerated Mixed-Precision Graph Optimization Framework
- User-Centric Communication Service Provision for Edge-Assisted Mobile Augmented Reality
- Advancing Data Quality of Marine Archaeological Documentation Using Underwater Robotics: From Simulation Environments to Real-World Scenarios
- Good Weights: Proactive, Adaptive Dead Reckoning Fusion for Continuous and Robust Visual SLAM
- Real-Time Indoor Object SLAM with LLM-Enhanced Priors
- PL-VIWO2: A Lightweight, Fast and Robust Visual-Inertial-Wheel Odometry Using Points and Lines
- MASt3R-Fusion: Integrating Feed-Forward Visual Model with IMU, GNSS for High-Functionality SLAM
- Decision-Driven Semantic Object Exploration for Legged Robots via Confidence-Calibrated Perception and Topological Subgoal Selection
- Deep learning-based 3D morphological segmentation and quantitative growth analysis of field-grown cabbage across the full cycle
- SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization
- Crater Observing Bio-inspired Rolling Articulator (COBRA)
- MV-UMI: A Scalable Multi-View Interface for Cross-Embodiment Learning
- HUNT: High-Speed UAV Navigation and Tracking in Unstructured Environments via Instantaneous Relative Frames
- Combining Projected Uncertainty for Self-Supervised Visual Odometry: From Two-Frame to Multi-Frame
- ProDyG: Progressive Dynamic Scene Reconstruction via Gaussian Splatting from Monocular Videos
- SLAM-Former: Putting SLAM into One Transformer
- See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
- Omni-LIVO: Robust RGB-Colored Multi-Camera Visual-Inertial-LiDAR Odometry via Photometric Migration and ESIKF Fusion
- BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots
- MCGS-SLAM: A Multi-Camera SLAM Framework Using Gaussian Splatting for High-Fidelity Mapping
- BIM Informed Visual SLAM for Construction Monitoring
- FlightDiffusion: Revolutionising Autonomous Drone Training with Diffusion Models Generating FPV Video
- ROOM: A Physics-Based Continuum Robot Simulator for Photorealistic Medical Datasets Generation
- Maps for Autonomous Driving: Full-process Survey and Frontiers
- VROOM - Visual Reconstruction over Onboard Multiview
- MemGS: Memory-Efficient Gaussian Splatting for Real-Time SLAM
- FastTrack: GPU-Accelerated Tracking for Visual SLAM
- Efficient and Accurate Downfacing Visual Inertial Odometry
- SMapper: A Multi-Modal Data Acquisition Platform for SLAM Benchmarking
- Online Dynamic SLAM with Incremental Smoothing and Mapping
- Radar-Based Odometry for Low-Speed Driving
- Real-time Photorealistic Mapping for Situational Awareness in Robot Teleoperation
- Evaluating Magic Leap 2 Tool Tracking for AR Sensor Guidance in Industrial Inspections
- IL-SLAM: Intelligent Line-assisted SLAM Based on Feature Awareness for Dynamic Environments
- Articulated Object Estimation in the Wild
- ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association
- SR-SLAM: Scene-reliability Based RGB-D SLAM in Diverse Environments
- FGO-SLAM: Enhancing Gaussian SLAM with Globally Consistent Opacity Radiance Field
- DyPho-SLAM : Real-time Photorealistic SLAM in Dynamic Environments
- Tank dataset: An underwater multi-sensor dataset for SLAM evaluation
- The Rosario Dataset v2: Multimodal Dataset for Agricultural Robotics
- UnPose: Uncertainty-Guided Diffusion Priors for Zero-Shot Pose Estimation
- HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation
- Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey
- EAROL: Environmental Augmented Perception-Aware Planning and Robust Odometry via Downward-Mounted Tilted LiDAR
- ROVER: Robust Loop Closure Verification with Trajectory Prior in Repetitive Environments
- Unifying Scale-Aware Depth Prediction and Perceptual Priors for Monocular Endoscope Pose Estimation and Tissue Reconstruction
- ViPE: Video Pose Engine for 3D Geometric Perception
- QoE-Aware Service Provision for Mobile AR Rendering: An Agent-Driven Approach
- XR Reality Check: What Commercial Devices Deliver for Spatial Tracking
- False Reality: Uncovering Sensor-induced Human-VR Interaction Vulnerability
- Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction
- Autonomous Navigation of Cloud-Controlled Quadcopters in Confined Spaces Using Multi-Modal Perception and LLM-Driven High Semantic Reasoning
- Bio-Inspired Topological Autonomous Navigation with Active Inference in Robotics
- Navigation and Exploration with Active Inference: from Biology to Industry
- EGS-SLAM: RGB-D Gaussian Splatting SLAM with Events
- RiemanLine: Riemannian Manifold Representation of 3D Lines for Factor Graph Optimization
- Opti-Acoustic Scene Reconstruction in Highly Turbid Underwater Environments
- SGAD: Semantic and Geometric-aware Descriptor for Local Feature Matching
- LT-Gaussian: Long-Term Map Update Using 3D Gaussian Splatting for Autonomous Driving
- CoProU-VO: Combining Projected Uncertainty for End-to-End Unsupervised Monocular Visual Odometry
- The Monado SLAM Dataset for Egocentric Visual-Inertial Tracking
- Stereo 3D Gaussian Splatting SLAM for Outdoor Urban Scenes
- DuLoc: Life-Long Dual-Layer Localization in Changing and Dynamic Expansive Scenarios
- Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
- Adaptive Prior Scene-Object SLAM for Dynamic Environments
- FMimic: Foundation Models are Fine-grained Action Learners from Human Videos
- FROSS: Faster-than-Real-Time Online 3D Semantic Scene Graph Generation from RGB-D Images
- DOA: A Degeneracy Optimization Agent with Adaptive Pose Compensation Capability based on Deep Reinforcement Learning
- G2S-ICP SLAM: Geometry-aware Gaussian Splatting ICP SLAM
- FAST-LIO2: Fast Direct LiDAR-Inertial Odometry
- 3D LiDAR SLAM: A survey
- MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware
- RemixFusion: Residual-based Mixed Representation for Large-scale Online RGB-D Reconstruction
- A Target-based Multi-LiDAR Multi-Camera Extrinsic Calibration System
- VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
- FFI-VTR: Lightweight and Robust Visual Teach and Repeat Navigation based on Feature Flow Indicator and Probabilistic Motion Planning
- DINO-VO: A Feature-based Visual Odometry Leveraging a Visual Foundation Model
- MoCap2GT: A High-Precision Ground Truth Estimator for SLAM Benchmarking Based on Motion Capture and IMU Fusion
- NemeSys: Toward Online Underwater Exploration with Remote Operator-in-the-loop Adaptive Autonomy
- CorrMoE: Mixture of Experts with De-stylization Learning for Cross-Scene and Cross-Domain Correspondence Pruning
- A Probability-guided Sampler for Neural Implicit Surface Rendering
- Princeton365: A Diverse Dataset with Accurate Camera Pose
- Comparison of Localization Algorithms between Reduced-Scale and Real-Sized Vehicles Using Visual and Inertial Sensors
- Multi-IMU Sensor Fusion for Legged Robots
- Towards Robust Sensor-Fusion Ground SLAM: A Comprehensive Benchmark and A Resilient Framework
- IRAF-SLAM: An Illumination-Robust and Adaptive Feature-Culling Front-End for Visual SLAM in Challenging Environments
- PacGDC: Label-Efficient Generalizable Depth Completion with Projection Ambiguity and Consistency
- g2o vs. Ceres: Optimizing Scan Matching in Cartographer SLAM
- Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion
- Piggyback Camera: Easy-to-Deploy Visual Surveillance by Mobile Sensing on Commercial Robot Vacuums
- VISC: mmWave Radar Scene Flow Estimation using Pervasive Visual-Inertial Supervision
- Gaussian-LIC2: LiDAR-Inertial-Camera Gaussian Splatting SLAM
- Query-Based Adaptive Aggregation for Multi-Dataset Joint Training Toward Universal Visual Place Recognition
- Outdoor Monocular SLAM with Global Scale-Consistent 3D Gaussian Pointmaps
- MGSfM: Multi-Camera Geometry Driven Global Structure-from-Motion
- The Geometric Observability Index: Influence, Fisher Information, and Weak Observability in \SE Pose Estimation
- Vision-based Goal-Reaching Control for Mobile Robots Using a Hierarchical Learning Framework
- Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction
- TVG-SLAM: Robust Gaussian Splatting SLAM with Tri-view Geometric Constraints
- Event-based Stereo Visual-Inertial Odometry with Voxel Map
- SPICE-HL3: Single-Photon, Inertial, and Stereo Camera dataset for Exploration of High-Latitude Lunar Landscapes
- EndoFlow-SLAM: Real-Time Endoscopic SLAM with Flow-Constrained Gaussian Splatting
- ZeroVO: Visual Odometry with Minimal Assumptions
- GRAND-SLAM: Local Optimization for Globally Consistent Large-Scale Multi-Agent Gaussian SLAM
- ThermalLoc: A Vision Transformer-Based Approach for Robust Thermal Camera Relocalization in Large-Scale Environments
- MCN-SLAM: Multi-Agent Collaborative Neural SLAM with Hybrid Implicit Neural Scene Representation
- Multimodal Fusion SLAM with Fourier Attention
- DualTHOR: A Dual-Arm Humanoid Simulation Platform for Contingency-Aware Planning
- ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models
- VIMS: A Visual-Inertial-Magnetic-Sonar SLAM System in Underwater Environments
- ChartBlender: An Interactive System for Authoring and Synchronizing Visualization Charts in Video
- Faster than Fast: Accelerating Oriented FAST Feature Detection on Low-end Embedded GPUs
- A Novel ViDAR Device With Visual Inertial Encoder Odometry and Reinforcement Learning-Based Active SLAM Method
- SuperPoint-SLAM3: Augmenting ORB-SLAM3 with Deep Features, Adaptive NMS, and Learning-Based Loop Closure
- Gaussian Mapping for Evolving Scenes
- Advances on Affordable Hardware Platforms for Human Demonstration Acquisition in Agricultural Applications
- Dy3DGS-SLAM: Monocular 3D Gaussian Splatting SLAM for Dynamic Environments
- GS4: Generalizable Sparse Splatting Semantic SLAM
- Enhancing Situational Awareness in Underwater Robotics with Multi-modal Spatial Perception
- On-the-fly Reconstruction for Large-Scale Novel View Synthesis from Unposed Images
- Learning to Drive Anywhere with Model-Based Reannotation
- cuVSLAM: CUDA accelerated visual odometry and mapping
- Dynamics and Control of Vision-Aided Multi-UAV-tethered Netted System Capturing Non-Cooperative Target
- GeneA-SLAM2: Dynamic SLAM with AutoEncoder-Preprocessed Genetic Keypoints Resampling and Depth Variance-Guided Dynamic Region Removal
- SEMNAV: A Semantic Segmentation-Driven Approach to Visual Semantic Navigation
- ADEPT: Adaptive Diffusion Environment for Policy Transfer Sim-to-Real
- Globally Consistent RGB-D SLAM with 2D Gaussian Splatting
- Understanding while Exploring: Semantics-driven Active Mapping
- UP-SLAM: Adaptively Structured Gaussian SLAM with Uncertainty Prediction in Dynamic Environments
- Sewer defect instance segmentation, localization, and 3D reconstruction for sewer floating capsule robots
- HS-SLAM: A Fast and Hybrid Strategy-Based SLAM Approach for Low-Speed Autonomous Driving
- CAD-SLAM: Consistency-Aware Dynamic SLAM with Dynamic-Static Decoupled Mapping
- CHOW-SLAM: Compact Hybrid Representation with Complementary Overlap Window Optimization for RGB-D SLAM
- Why Not Replace? Sustaining Long-Term Visual Localization via Handcrafted-Learned Feature Collaboration on CPU
- UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization
- TANGO-VIO: Triangulation-Aware Navigation with Guaranteed Feature-Observability for Visual-Inertial Odometry
- RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation
- Stipple: Real-Time Incremental Gaussian Splatting with Visual-Inertial Tracking
- VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
- Is Semantic SLAM Ready for Embedded Systems ? A Comparative Survey
- Structureless VIO
- Look Up and Look Back: Hidden Attention and Latent Orientation in a Frozen Foundation Model for Panoramic SLAM
- TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation
- Large-Scale Gaussian Splatting SLAM
- Edge-Enabled VIO with Long-Tracked Features for High-Accuracy Low-Altitude IoT Navigation
- SafeNav: Safe Path Navigation using Landmark Based Localization in a GPS-denied Environment
- Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes
- STEMbot: A Compliant Robot for Under-Canopy Plant Navigation
- Disturbance-aware Motion Planning for Over-actuated Underwater Vehicles Exploiting Actuation Redundancy for High-fidelity 3D Reconstruction
- OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer
- Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
- Efficient Dense Matching for Enhanced Gaussian Splatting Using AV1 Motion Vectors
- LPVIMO-SAM: Tightly-coupled LiDAR/Polarization Vision/Inertial/Magnetometer/Optical Flow Odometry via Smoothing and Mapping
- GSFeatLoc: Visual Localization Using Feature Correspondence on 3D Gaussian Splatting
- ReefMapGS: Enabling Large-Scale Underwater Reconstruction by Closing the Loop Between Multimodal SLAM and Gaussian Splatting
- Keyframe-Based Feed-Forward Visual Odometry
- VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction
- Fauna Sprout: A lightweight, approachable, developer-ready humanoid robot
- Screw-based feature constraint model and degeneracy analysis for robotic state estimation: Theory and experiments
- Certifiably-Correct Mapping for Safe Navigation Despite Odometry Drift
- EdgePoint2: Compact Descriptors for Superior Efficiency and Accuracy
- Bias-Eliminated PnP for Stereo Visual Odometry: Provably Consistent and Large-Scale Localization
- Differential 6-DOF Pose Estimation with Provable First-Order Immunity to Camera Calibration Errors
- ToF-Splatting: Dense SLAM using Sparse Time-of-Flight Depth and Multi-Frame Integration
- Perception and Navigation in Autonomous Systems in the Era of Learning: A Survey
- SG-Reg: Generalizable and Efficient Scene Graph Registration
- Unreal Robotics Lab: A High-Fidelity Robotics Simulator with Advanced Physics and Rendering
- SLAM&Render: A Benchmark for the Intersection Between Neural Rendering, Gaussian Splatting and SLAM
- On Incremental Structure from Motion Using Lines
- An Online Adaptation Method for Robust Depth Estimation and Visual Odometry in the Open World
- Nonlinear Observer Design for Landmark-Inertial Simultaneous Localization and Mapping
- JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment
- SA-LIVO: Efficient LiDAR-Inertial-Visual Odometry with Subspace-Aware Degeneracy Handling
- SeeTree -- A modular, open-source system for tree detection and orchard localization
- DSM: Constructing a Diverse Semantic Map for 3D Visual Grounding
- Knowledge Distillation for Underwater Feature Extraction and Matching via GAN-synthesized Images
- UWB Anchor Based Localization of a Planetary Rover
- D2USt3R: Enhancing 3D Reconstruction for Dynamic Scenes
- Embracing Dynamics: Dynamics-aware 4D Gaussian Splatting SLAM
Related