DUSt3R: Geometric 3D Vision Made Easy
2023/12/21 by Shuzhe Wang, Wang, Shuzhe, Vincent Leroy +7 · 576 citations
Computer Science · Engineering · #Advanced Vision and Imaging #Optical measurement and interference techniques #Image Processing Techniques and Applications
paper · pdf · doi:10.48550/arxiv.2312.14132
Abstract
Multi-view stereo reconstruction (MVS) in the wild requires to first estimate the camera parameters e.g. intrinsic and extrinsic parameters. These are usually tedious and cumbersome to obtain, yet they are mandatory to triangulate corresponding pixels in 3D space, which is the core of all best performing MVS algorithms. In this work, we take an opposite stance and introduce DUSt3R, a radically novel paradigm for Dense and Unconstrained Stereo 3D Reconstruction of arbitrary image collections, i.e. operating without prior information about camera calibration nor viewpoint poses. We cast the pairwise reconstruction problem as a regression of pointmaps, relaxing the hard constraints of usual projective camera models. We show that this formulation smoothly unifies the monocular and binocular reconstruction cases. In the case where more than two images are provided, we further propose a simple yet effective global alignment strategy that expresses all pairwise pointmaps in a common reference frame. We base our network architecture on standard Transformer encoders and decoders, allowing us to leverage powerful pretrained models. Our formulation directly provides a 3D model of the scene as well as depth information, but interestingly, we can seamlessly recover from it, pixel matches, relative and absolute camera. Exhaustive experiments on all these tasks showcase that the proposed DUSt3R can unify various 3D vision tasks and set new SoTAs on monocular/multi-view depth estimation as well as relative pose estimation. In summary, DUSt3R makes many geometric 3D vision tasks easy.
Cited by
- Stochastic Siamese MAE Pretraining for Longitudinal Medical Images
- RealX3D: A Physically-Degraded 3D Benchmark for Multi-view Visual Restoration and Reconstruction
- SpatialMosaic: A Multiview VLM Dataset for Partial Visibility
- 3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
- GenSplatCodec: Feed-Forward Gaussian Splatting Compression via One-Step Diffusion
- KV-Tracker: Real-Time Pose Tracking with Transformers
- Long-tail Internet photo reconstruction
- Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation
- Compressing Observation History into Agent Memory: Distilling Transformers into Recurrent Transformers
- RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction
- Calibration-Free 3D Multi-Camera People Tracking for Indoor Environment
- Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video
- FUSE-Flow: A Decoupled Framework for Calibration and Stateless Real-Time Multi-View Point Cloud Fusion
- ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing
- Geometric Context Transformer for Streaming 3D Reconstruction
- Sky2Ground: A Benchmark for Site Modeling under Varying Altitude
- ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
- Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
- Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
- Analyzing the Mechanism of Attention Collapse in VGGT from a Dynamics Perspective
- How Much 3D Do Video Foundation Models Encode?
- Pix2NPHM: Learning to Regress NPHM Reconstructions From a Single Image
- WorldWarp: Propagating 3D Geometry with Asynchronous Video Diffusion
- Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface
- ReDepth Anything: Test-Time Depth Refinement via Self-Supervised Re-lighting
- Dexterous World Models
- Robust Scene Coordinate Regression via Geometrically-Consistent Global Descriptors
- G3Splat: Geometrically Consistent Generalizable Gaussian Splatting
- FLEG: Feed-Forward Language Embedded Gaussian Splatting from Any Views via Compact Semantic Representation
- DVGT: Driving Visual Geometry Transformer
- SceneDiff: A Benchmark and Method for Multiview Object Change Detection
- Make-It-Poseable: Feed-forward Latent Posing Model for 3D Characters
- 4D Primitive-Mâché: Glueing Primitives for Persistent 4D Scene Reconstruction
- Privacy-Aware Sharing of Raw Spatial Sensor Data for Cooperative Perception
- In Pursuit of Pixel Supervision for Visual Pre-training
- Robust Multi-view Camera Calibration from Dense Matches
- MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors
- EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligence
- Spatia: Video Generation with Updatable Spatial Memory
- Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian Splatting
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction
- Unified Semantic Transformer for 3D Scene Understanding
- LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
- Charge: A Comprehensive Novel View Synthesis Benchmark and Dataset to Bind Them All
- DePT3R: Joint Dense Point Tracking and 3D Reconstruction of Dynamic Scenes in a Single Forward Pass
- Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
- Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
- On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
- VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation
- SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
- Any4D: Unified Feed-Forward Metric 4D Reconstruction
- Sharp Monocular View Synthesis in Less Than a Second
- PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View Reasoning
- GLaD: Geometric Latent Distillation for Vision-Language-Action Models
- StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
- GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
- Efficiently Reconstructing Dynamic Scenes One D4RT at a Time
- OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics
- TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels
- Inferring Compositional 4D Scenes without Ever Seeing One
- Trajectory Densification and Depth from Perspective-based Blur
- Selfi: Self Improving Reconstruction Engine via 3D Geometric Feature Alignment
- MuSASplat: Efficient Sparse-View 3D Gaussian Splats via Lightweight Multi-Scale Adaptation
- WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- Dynamic Visual SLAM using a General 3D Prior
- HuPrior3R: Incorporating Human Priors for Better 3D Dynamic Reconstruction from Monocular Videos
- The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
- 4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
- LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
- Splat-SAP: Feed-Forward Gaussian Splatting for Human-Centered Scene with Scale-Aware Point Map Reconstruction
- ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
- Towards Cross-View Point Correspondence in Vision-Language Models
- Unique Lives, Shared World: Learning from Single-Life Videos
- C3G: Learning Compact 3D Representations with 2K Gaussians
- ReCamDriving: LiDAR-Free Camera-Controlled Video Synthesis for Novel Trajectories
- What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
- Emergent Outlier View Rejection in Visual Geometry Grounded Transformers
- MUT3R: Motion-aware Updating Transformer for Dynamic 3D Reconstruction
- Flux4D: Flow-based Unsupervised 4D Reconstruction
- CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
- Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
- AVGGT: Rethinking Global Attention for Accelerating VGGT
- Taming Camera-Controlled Video Generation with Verifiable Geometry Reward
- Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
- MV-TAP: Tracking Any Point in Multi-View Videos
- KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM
- Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
- S2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
- EAG3R: Event-Augmented 3D Geometry Estimation for Dynamic and Extreme-Lighting Scenes
- 3D-Consistent Multi-View Editing by Correspondence Guidance
- Asset-Driven Sematic Reconstruction of Dynamic Scene with Multi-Human-Object Interactions
- TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completion
- Cross-Temporal 3D Gaussian Splatting for Sparse-View Guided Scene Update
- GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence
- GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation
- SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
- VG3T: Visual Geometry Grounded Gaussian Transformer
- AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
- GSpaRC: Gaussian Splatting for Real-time Reconstruction of RF Channels
- Emergent Extreme-View Geometry in 3D Foundation Models
- Gaussians on Fire: High-Frequency Reconstruction of Flames
- Combining Projected Uncertainty for Self-Supervised Visual Odometry: From Two-Frame to Multi-Frame
- Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge Distillation
- DriveVGGT: Calibration-Constrained Visual Geometry Transformers for Multi-Camera Autonomous Driving
- ColonAdapter: Geometry Estimation Through Foundation Model Adaptation for Colonoscopy
- Geometry Meets Light: Leveraging Geometric Priors for Universal Photometric Stereo under Limited Multi-Illumination Cues
- HTTM: Head-wise Temporal Token Merging for Faster VGGT
- CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
- G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
- Vision-Language Memory for Spatial Reasoning
- Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
- PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
- Zoo3D: Zero-Shot 3D Object Detection at Scene Level
- VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend
- Multi-Agent Monocular Dense SLAM With 3D Reconstruction Priors
- View-Consistent Diffusion Representations for 3D-Consistent Video Generation
- One4D: Unified 4D Generation and Reconstruction via Decoupled LoRA Control
- Sphinx: Efficiently Serving Novel View Synthesis using Regression-Guided Selective Refinement
- 4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
- SwiftVGGT: A Scalable Visual Geometry Grounded Transformer for Large-Scale Scenes
- C3Po: Cross-View Cross-Modality Correspondence by Pointmap Prediction
- Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
- Novel View Synthesis from A Few Glimpses via Test-Time Natural Video Completion
- SPIDER: Spatial Image CorresponDence Estimator for Robust Calibration
- HALO: High-Altitude Language-Conditioned Monocular Aerial Exploration and Navigation
- NoPe-NeRF++: Local-to-Global Optimization of NeRF with No Pose Prior
- SING3R-SLAM: Submap-based Indoor Monocular Gaussian SLAM with 3D Reconstruction Priors
- DepthFocus: Controllable Depth Estimation for See-Through Scenes
- SVRecon: Sparse Voxel Rasterization for Surface Reconstruction
- NoPo-Avatar: Generalizable and Animatable Avatars from Sparse Inputs without Human Poses
- POMA-3D: The Point Map Way to 3D Scene Understanding
- CylinderDepth: Cylindrical Spatial Attention for Multi-View Consistent Self-Supervised Surround Depth Estimation
- PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day Photos
- ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy Estimation
- SceneEdited: A City-Scale Benchmark for 3D HD Map Updating via Image-Guided Change Detection
- RoMa v2: Harder Better Faster Denser Feature Matching
- InstructMix2Mix: Consistent Sparse-View Editing Through Multi-View Model Personalization
- Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers
- BEDLAM2.0: Synthetic Humans and Cameras in Motion
- Dental3R: Geometry-Aware Pairing for Intraoral 3D Reconstruction from Sparse-View Photographs
- CloseUpShot: Close-up Novel View Synthesis from Sparse-views via Point-conditioned Diffusion Model
- CLIDD: Cross-Layer Independent Deformable Description for Efficient and Discriminative Local Feature Representation
- Lumos3D: A Single-Forward Framework for Low-Light 3D Scene Restoration
- MVSMamba: Multi-View Stereo with State Space Model
- EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision
- Sparse3DPR: Training-Free 3D Hierarchical Scene Parsing and Task-Adaptive Subgraph Reasoning from Sparse RGB Views
- Wid3R: Wide Field-of-View 3D Reconstruction via Camera Model Conditioning
- S-MUSt3R: Sliding Multi-view 3D Reconstruction
- YoNoSplat: You Only Need One Model for Feedforward 3D Gaussian Splatting
- 4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic Scenes from Monocular Videos
- UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene Reconstruction
- GraspView: Active Perception Scoring and Best-View Optimization for Robotic Grasping in Cluttered Environments
- Room Envelopes: A Synthetic Dataset for Indoor Layout Reconstruction from Images
- DentalSplat: Dental Occlusion Novel View Synthesis from Sparse Intra-Oral Photographs
- LiDAR-VGGT: Cross-Modal Coarse-to-Fine Fusion for Globally Consistent and Metric-Scale Dense Mapping
- DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
- Densemarks: Learning Canonical Embeddings for Human Heads Images via Point Tracks
- GauDP: Reinventing Multi-Agent Collaboration through Gaussian-Image Synergy in Diffusion Policies
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
- SEE4D: Pose-Free 4D Generation via Auto-Regressive Video Inpainting
- Glob3R: Global Structure-from-Motion with 3D Foundation Models
- Surflo: Consistent 3D Surface Flow Model with Global State
- NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction
- UniForward: Unified 3D Scene and Semantic Field Reconstruction via Feed-Forward Gaussian Splatting from Only Sparse-View Images
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
- MoonAnything: A Vision Benchmark with Large-Scale Lunar Supervised Data
- Understanding and Optimizing Attention-Based Sparse Matching for Diverse Local Features
- XRefine: Attention-Guided Keypoint Match Refinement
- SpaceTimePilot: Generative Rendering of Dynamic Scenes Across Space and Time
- Understanding Multi-View Transformers
- Kineo: Calibration-Free Metric Motion Capture From Sparse RGB Cameras
- PlanarGS: High-Fidelity Indoor 3D Gaussian Splatting Guided by Vision-Language Planar Priors
- Adaptive Keyframe Selection for Scalable 3D Scene Reconstruction in Dynamic Environments
- ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation
- IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
- CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling
- OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects
- PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning
- AnyPcc: Compressing Any Point Cloud with a Single Universal Model
- VGD: Visual Geometry Gaussian Splatting for Feed-Forward Surround-view Driving Reconstruction
- PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
- Advances in 4D Representation: Geometry, Motion, and Interaction
- Robust Driving QA through Metadata-Grounded Context and Task-Specific Prompts
- GeoDiff: Geometry-Guided Diffusion for Metric Depth Estimation
- PFGS: Pose-Fused 3D Gaussian Splatting for Complete Multi-Pose Object Reconstruction
- DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
- PLANA3R: Zero-shot Metric Planar 3D Reconstruction via Feed-Forward Planar Splatting
- Adapting Stereo Vision From Objects To 3D Lunar Surface Reconstruction with the StereoLunar Dataset
- From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
- Pursuing Minimal Sufficiency in Spatial Reasoning
- SaLon3R: Structure-aware Long-term Generalizable 3D Reconstruction from Unposed Images
- Terra: Explorable Native 3D World Model with Point Latents
- C4D: 4D Made from 3D through Dual Correspondences
- GOPLA: Generalizable Object Placement Learning via Synthetic Augmentation of Human Arrangement
- SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation
- MatchAttention: Matching the Relative Positions for High-Resolution Cross-View Matching
- Learning Neural Parametric 3D Breast Shape Models for Metrical Surface Reconstruction From Monocular RGB Videos
- Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
- DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
- Trace Anything: Representing Any Video in 4D via Trajectory Fields
- Scene Coordinate Reconstruction Priors
- G4Splat: Geometry-Guided Gaussian Splatting with Generative Prior
- ACE-G: Improving Generalization of Scene Coordinate Regression Through Query Pre-Training
- FVO: Fast Visual Odometry with Transformers
- WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting
- AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes
- EC3R-SLAM: Efficient and Consistent Monocular Dense SLAM with Feed-Forward 3D Reconstruction
- Gesplat: Robust Pose-Free 3D Reconstruction via Geometry-Guided Gaussian Splatting
- Geometry-Aware Scene Configurations for Novel View Synthesis
- Online Video Depth Anything: Temporally-Consistent Depth Prediction with Low Memory Consumption
- SViM3D: Stable Video Material Diffusion for Single Image 3D Generation
- XYZCylinder: Towards Compatible Feed-Forward 3D Gaussian Splatting for Driving Scenes via Unified Cylinder Lifting Method
- ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation
- DynamicEval: Rethinking Evaluation for Dynamic Text-to-Video Synthesis
- MoRe: Monocular Geometry Refinement via Graph Optimization for Cross-View Consistency
- Human3R: Everyone Everywhere All at Once
- Scaling Sequence-to-Sequence Generative Neural Rendering
- Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft
- FSFSplatter: Build Surface and Novel Views with Sparse-Views within 2min
- EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory
- Instant4D: 4D Gaussian Splatting in Minutes
- A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features
- HART: Human Aligned Reconstruction Transformer
- Stylos: Multi-View 3D Stylization with Single-Forward Gaussian Splatting
- TTT3R: 3D Reconstruction as Test-Time Training
- GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification
- LVT: Large-Scale Scene Reconstruction via Local View Transformers
- Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning Synergy
- PROFusion: Robust and Accurate Dense Reconstruction via Camera Pose Regression and Optimization
- Forge4D: Feed-Forward 4D Human Reconstruction and Interpolation from Uncalibrated Sparse-view Videos
- Plant3R: Fusing 3D feature learning with Gaussian splatting to enhance wheat plant 3D reconstruction precision
- VGGT-X: When VGGT Meets Dense Novel View Synthesis
- UniLat3D: Geometry-Appearance Unified Latents for Single-Stage 3D Generation
- BOSfM: A View Planning Framework for Optimal 3D Reconstruction of Agricultural Scenes
- RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph Optimization
- GRS-SLAM3R: Real-Time Dense SLAM with Gated Recurrent State
- ZeroScene: A Zero-Shot Framework for 3D Scene Generation from a Single Image and Controllable Texture Editing
- From Fields to Splats: A Cross-Domain Survey of Real-Time Neural Scene Representations
- OVSeg3R: Learn Open-vocabulary Instance Segmentation from 2D via 3D Reconstruction
- DiffTex: Differentiable Texturing for Architectural Proxy Models
- OracleGS: Grounding Generative Priors for Sparse-View Gaussian Splatting
- GeLoc3r: Enhancing Relative Camera Pose Regression with Geometric Consistency Regularization
- FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction
- Quantized Visual Geometry Grounded Transformer
- MASt3R-Fusion: Integrating Feed-Forward Visual Model with IMU, GNSS for High-Functionality SLAM
- Efficient Construction of Implicit Surface Models From a Single Image for Motion Generation
- Dense Semantic Matching with VGGT Prior
- Joint Flow Trajectory Optimization For Feasible Robot Motion Generation from Video Demonstrations
- MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation
- Cross-Modal Instructions for Robot Motion Generation
- Reflect3r: Single-View 3D Stereo Reconstruction Aided by Mirror Reflections
- 4D Driving Scene Generation With Stereo Forcing
- TUN3D: Towards Real-World Scene Understanding from Unposed Images
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
- 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image
- 3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
- LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
- Dark3R: Learning Structure from Motion in the Dark
- DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation
- SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation
- DeblurSplat: SfM-free 3D Gaussian Splatting with Event Camera for Robust Deblurring
- VGGT-DP: Generalizable Robot Control via Vision Foundation Models
- VideoArtGS: Building Digital Twins of Articulated Objects from Monocular Video
- SmokeSeer: 3D Gaussian Splatting for Smoke Removal and Scene Reconstruction
- Evict3R: Training-Free Token Eviction for Memory-Bounded Streaming Visual Geometry Transformers
- ProDyG: Progressive Dynamic Scene Reconstruction via Gaussian Splatting from Monocular Videos
- OrthoLoC: UAV 6-DoF Localization and Calibration Using Orthographic Geodata
- SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
- Efficient 3D Scene Reconstruction and Simulation from Sparse Endoscopic Views
- SLAM-Former: Putting SLAM into One Transformer
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
- IDU: Incremental Dynamic Update of Existing 3D Virtual Environments with New Imagery Data
- SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
- GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation
- MeshSplat: Generalizable Sparse-View Surface Reconstruction via Gaussian Splatting
- FingerSplat: Contactless Fingerprint 3D Reconstruction and Generation based on 3D Gaussian Splatting
- MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the Wild
- Lightweight and Accurate Multi-View Stereo with Confidence-Aware Diffusion Model
- RGB-Only Supervised Camera Parameter Optimization in Dynamic Scenes
- BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots
- Explicit Context-Driven Neural Acoustic Modeling for High-Fidelity RIR Generation
- Gaussian Alignment for Relative Camera Pose Estimation via Single-View Reconstruction
- GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
- ROOM: A Physics-Based Continuum Robot Simulator for Photorealistic Medical Datasets Generation
- 3D Aware Region Prompted Vision Language Model
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
- Segmentation-Driven Initialization for Sparse-view 3D Gaussian Splatting
- WildSmoke: Ready-to-Use Dynamic 3D Smoke Assets from a Single Video in the Wild
- Leveraging Geometric Priors for Unaligned Scene Change Detection
- Calib3R: A 3D Foundation Model for Multi-Camera to Robot Calibration and 3D Metric-Scaled Scene Reconstruction
- CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis
- Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
- FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion Bases
- WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
- PAOLI: Pose-free Articulated Object Learning from Sparse-view Images
- HOSt3R: Keypoint-free Hand-Object 3D Reconstruction from RGB images
- HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction
- FastVGGT: Training-Free Acceleration of Visual Geometry Transformer
- ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association
- Advances and Trends in the 3D Reconstruction of the Shape and Motion of Animals
- NeuralMeshing: Complete Object Mesh Extraction from Casual Captures
- PlayerOne: Egocentric World Simulator
- Complete Gaussian Splats from a Single Image with Denoising Diffusion Models
- Multi-View 3D Point Tracking
- SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass
- Adam SLAM - the last mile of camera calibration with 3DGS
- FastAvatar: Towards Unified Fast High-Fidelity 3D Avatar Reconstruction with Large Gaussian Reconstruction Transformers
- DriveSplat: Decoupled Driving Scene Reconstruction with Geometry-enhanced Partitioned Neural Gaussians
- Enhancing Novel View Synthesis from extremely sparse views with SfM-free 3D Gaussian Splatting Framework
- Snap-Snap: Taking Two Images to Reconstruct 3D Human Gaussians in Milliseconds
- LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos
- Sketch3DVE: Sketch-based 3D-Aware Scene Video Editing
- ROVER: Robust Loop Closure Verification with Trajectory Prior in Repetitive Environments
- 4DNeX: Feed-Forward 4D Generative Modeling Made Easy
- Quantifying and Alleviating Co-Adaptation in Sparse-View 3D Gaussian Splatting
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- Towards Understanding 3D Vision: the Role of Gaussian Curvature
- G-CUT3R: Guided 3D Reconstruction with Camera and Depth Prior Integration
- STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
- ViewBridge:Revisiting Cross-View Localization from Image Matching
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- ViPE: Video Pose Engine for 3D Geometric Perception
- Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction
- Dream4D: Lifting Camera-Controlled I2V towards Spatiotemporally Consistent 4D Generation
- RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
- Uncertainty Quantification Framework for Aerial and UAV Photogrammetry through Error Propagation
- FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion Distillation
- Multi-view Gaze Target Estimation
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- UGOD: Uncertainty-Guided Differentiable Opacity and Soft Dropout for Enhanced Sparse-View 3DGS
- Pseudo Depth Meets Gaussian: A Feed-forward RGB SLAM Baseline
- Surf3R: Rapid Surface Reconstruction from Sparse RGB Views in Seconds
- PIS3R: Very Large Parallax Image Stitching via Deep 3D Reconstruction
- IDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Control
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images
- DisCo3D: Distilling Multi-View Consistency for 3D Scene Editing
- No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
- DELTAv2: Accelerating Dense 3D Tracking
- Hestia: Voxel-Face-Aware Hierarchical Next-Best-View Acquisition for Efficient 3D Reconstruction
- GECO: Geometrically Consistent Embedding with Lightspeed Inference
- CoProU-VO: Combining Projected Uncertainty for End-to-End Unsupervised Monocular Visual Odometry
- MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion
- π3: Permutation-Equivariant Visual Geometry Learning
- iLRM: An Iterative Large 3D Reconstruction Model
- UFV-Splatter: Pose-Free Feed-Forward 3D Gaussian Splatting Adapted to Unfavorable Views
- Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
- PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction
- Reconstructing 4D Spatial Intelligence: A Survey
- GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting
- LONG3R: Long Sequence Streaming 3D Reconstruction
- Unposed 3DGS Reconstruction with Probabilistic Procrustes Mapping
- Stereo-GS: Multi-View Stereo Vision Model for Generalizable 3D Gaussian Splatting Reconstruction
- Towards Geometric and Textural Consistency 3D Scene Generation via Single Image-guided Model Generation and Layout Optimization
- An Evaluation of DUSt3R/MASt3R/VGGT 3D Reconstruction on Photogrammetric Aerial Blocks
- Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey
- InvRGB+L: Inverse Rendering of Complex Scenes with Unified Color and LiDAR Reflectance Modeling
- VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
- Sparse-View 3D Reconstruction: Recent Advances and Open Challenges
- Dens3R: A Foundation Model for 3D Geometry Prediction
- LongSplat: Online Generalizable 3D Gaussian Splatting from Long Sequence Images
- The Less You Depend, The More You Learn: Synthesizing Novel Views from Sparse, Unposed Images with Minimal 3D Knowledge
- SpatialTrackerV2: 3D Point Tracking Made Easy
- BRUM: Robust 3D Vehicle Reconstruction from 360 Sparse Images
- Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models
- Princeton365: A Diverse Dataset with Accurate Camera Pose
- StreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video Streams
- iTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos
- TRAN-D: 2D Gaussian Splatting-based Sparse-view Transparent Object Depth Reconstruction via Physics Simulation for Scene Update
- Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation
- Streaming 4D Visual Geometry Transformer
- Kaleidoscopic Background Attack: Disrupting Pose Estimation with Multi-Fold Radial Symmetry Textures
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second
- Review of Feed-forward 3D Reconstruction: From DUSt3R to VGGT
- RegGS: Unposed Sparse Views Gaussian Splatting with 3DGS Registration
- PacGDC: Label-Efficient Generalizable Depth Completion with Projection Ambiguity and Consistency
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
- Piggyback Camera: Easy-to-Deploy Visual Surveillance by Mobile Sensing on Commercial Robot Vacuums
- RIPE: Reinforcement Learning on Unlabeled Image Pairs for Robust Keypoint Extraction
- Learning Robust Stereo Matching in the Wild with Selective Mixture-of-Experts
- Voyaging into Perpetual Dynamic Scenes from a Single View
- Outdoor Monocular SLAM with Global Scale-Consistent 3D Gaussian Pointmaps
- Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory
- LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
- SIU3R: Simultaneous Scene Understanding and 3D Reconstruction Beyond Feature Alignment
- MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details
- HyperGaussians: High-Dimensional Gaussian Splatting for High-Fidelity Animatable Face Avatars
- I3DM: Implicit 3D-aware Memory Retrieval and Injection for Consistent Video Scene Generation
- DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation
- Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
- UFM: A Simple Path towards Unified Dense Correspondence with Flow
- Geometry-aware 4D Video Generation for Robot Manipulation
- Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction
- AttentionGS: Towards Initialization-Free 3D Gaussian Splatting via Structural Attention
- Jamais Vu: Exposing the Generalization Gap in Supervised Semantic Correspondence
- GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra
- TVG-SLAM: Robust Gaussian Splatting SLAM with Tri-view Geometric Constraints
- AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D Generation
- EgoM2P: Egocentric Multimodal Multitask Pretraining
- RoboPearls: Editable Video Simulation for Robot Manipulation
- Video Perception Models for 3D Scene Synthesis
- StereoDiff: Stereo-Diffusion Synergy for Video Depth Estimation
- WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration
- CoCo4D: Comprehensive and Complex 4D Scene Generation
- VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
- 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
- YouTube-Occ: Learning Indoor 3D Semantic Occupancy Prediction from YouTube Videos
- MCN-SLAM: Multi-Agent Collaborative Neural SLAM with Hybrid Implicit Neural Scene Representation
- BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
- Co-VisiON: Co-Visibility ReasONing on Sparse Image Sets of Indoor Scenes
- 4Real-Video-V2: Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation
- Semantic and Feature Guided Uncertainty Quantification of Visual Localization for Autonomous Vehicles
- RA-NeRF: Robust Neural Radiance Field Reconstruction with Accurate Camera Pose Estimation under Complex Trajectories
- BoxFusion: Reconstruction-Free Open-Vocabulary 3D Object Detection via Real-Time Multi-View Box Fusion
- SpatialLM: Training Large Language Models for Structured Indoor Modeling
- ProSplat: Improved Feed-Forward 3D Gaussian Splatting for Wide-Baseline Sparse Views
- LHM++: An Efficient Large Human Reconstruction Model for Pose-free Images to 3D
- Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
- SuperPlace: The Renaissance of Classical Feature Aggregation for Visual Place Recognition in the Era of Foundation Models
- Test3R: Learning to Reconstruct 3D at Test Time
- UNO: Unified Self-Supervised Monocular Odometry for Platform-Agnostic Deployment
- VEIGAR: View-consistent Explicit Inpainting and Geometry Alignment for 3D object Removal
- Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
- SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis
- GenWorld: Towards Detecting AI-generated Real-world Simulation Videos
- Splat and Replace: 3D Reconstruction with Repetitive Elements
- CryoFastAR: Fast Cryo-EM Ab Initio Reconstruction Made Easy
- GS4: Generalizable Sparse Splatting Semantic SLAM
- On-the-fly Reconstruction for Large-Scale Novel View Synthesis from Unposed Images
- Revisiting Depth Representations for Feed-Forward 3D Gaussian Splatting
- RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion
- Rectified Point Flow: Generic Point Cloud Pose Estimation
- OGGSplat: Open Gaussian Growing for Generalizable Reconstruction with Expanded Field-of-View
- Video World Models with Long-term Spatial Memory
- EX-4D: EXtreme Viewpoint 4D Video Synthesis via Depth Watertight Mesh
- Deep Learning Reforms Image Matching: A Survey and Outlook
- UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting
- Object-centric 3D Motion Field for Robot Learning from Human Videos
- FastMap: Revisiting Structure from Motion through First-Order Optimization
- Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
- PlückeRF: A Line-based 3D Representation for Few-view Reconstruction
- Generative Perception of Shape and Material from Differential Motion
- Towards In-the-wild 3D Plane Reconstruction from a Single Image
- E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
- SAB3R: Semantic-Augmented Backbone in 3D Reconstruction
- SteerPose: Simultaneous Extrinsic Camera Calibration and Matching from Articulation
- Rig3R: Rig-Aware Conditioning for Learned 3D Reconstruction
- DeepVerse: 4D Autoregressive Video Generation as a World Model
- Globally Consistent RGB-D SLAM with 2D Gaussian Splatting
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
- UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation
- Bi-Manual Joint Camera Calibration and Scene Representation
- AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views
- SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
- GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control
- EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
- ProBA: Probabilistic Bundle Adjustment with the Bhattacharyya Coefficient
- Styl3R: Instant 3D Stylized Reconstruction for Arbitrary Scenes and Styles
- Intern-GS: Vision Model Guided Sparse-View 3D Gaussian Splatting
- Sparse2DGS: Sparse-View Surface Reconstruction using 2D Gaussian Splatting with Dense Point Cloud
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
- Improving Novel view synthesis of 360^∘ Scenes in Extremely Sparse Views by Jointly Training Hemisphere Sampled Synthetic Images
- Sparfels: Fast Reconstruction from Sparse Unposed Imagery
- Focus What Matters: Matchability-Based Reweighting for Local Feature Matching
- From Single Images to Motion Policies via Video-Generation Environment Representations
- WonderPlay: Dynamic 3D Scene Generation from a Single Image and Actions
- SplatCo: Structure-View Collaborative Gaussian Splatting for Detail-Preserving Rendering of Large-Scale Unbounded Scenes
- Multi-View Unified Camera Fields: Geometry-Shaped Action-Facing Representations for RGB-Only Multi-Camera VLA Policies
- UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization
- SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
- X-GRM: Large Gaussian Reconstruction Model for Sparse-view X-rays to Computed Tomography
- Diving into the Fusion of Monocular Priors for Generalized Stereo Matching
- 3D Visual Illusion Depth Estimation
- Recollection from Pensieve: Novel View Synthesis via Learning from Uncalibrated Videos
- VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
- QuerySplat: Decoupling Geometry and Appearance Representations in 3DGS Prediction
- SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations
- MonoMobility: Zero-Shot 3D Mobility Analysis from Monocular Videos
- Look Up and Look Back: Hidden Attention and Latent Orientation in a Frozen Foundation Model for Panoramic SLAM
- Swimm3R: Splatting with Medium-aware SfM for Underwater 3D Reconstruction
- Depth Anything with Any Prior
- Advances in Radiance Field for Dynamic Scene: From Neural Field to Gaussian Field
- ExploreGS: a vision-based low overhead framework for 3D scene reconstruction
- Real2Render2Real: Scaling Robot Data Without Dynamics Simulation or Robot Hardware
- When Dance Video Archives Challenge Computer Vision
- AirSplat: Alignment and Rating for Robust Feed-Forward 3D Gaussian Splatting
- MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model
- 4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation
- DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion
- MVHumanNet++: A Large-scale Dataset of Multi-view Daily Dressing Human Captures with Richer Annotations for 3D Human Digitization
- AquaGS: Fast Underwater Scene Reconstruction with SfM-Free Gaussian Splatting
- Learning Multi-frame and Monocular Prior for Estimating Geometry in Dynamic Scenes
- StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision
- Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes
- Aerial Path Online Planning for Urban Scene Updation
- RayZer: A Self-supervised Large View Synthesis Model
- SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
- InCaRPose: In-Cabin Relative Camera Pose Estimation Model and Dataset
- Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction
- 2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction
- Quo Vadis, World Modeling?
- Vision as Unified Multimodal Generation
- VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
- OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer
- Spectral Probing of Feature Upsamplers in 2D-to-3D Scene Reconstruction
- Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
- GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
- IVGT: Implicit Visual Geometry Transformer for Neural Scene Representation
- VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction
- Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction
- Syn4D: A Multiview Synthetic 4D Dataset
- Pose-Aware Diffusion for 3D Generation
- Dance Style Recognition Using Laban Movement Analysis
- Large-scale visual SLAM for in-the-wild videos
- Generalizable Sparse-View 3D Reconstruction from Unconstrained Images
- RayTun3R: Online Camera Adaptation in 3D Foundation Models
- Remember to be Curious: Episodic Context and Persistent Worlds for 3D Exploration
- 3D-Belief: Embodied Belief Inference via Generative 3D World Modeling
- ReefMapGS: Enabling Large-Scale Underwater Reconstruction by Closing the Loop Between Multimodal SLAM and Gaussian Splatting
- UniRecGen: Unifying Multi-View 3D Reconstruction and Generation
- OccAny: Generalized Unconstrained Urban 3D Occupancy
- Marginalized Bundle Adjustment: Multi-View Camera Pose from Monocular Depth Estimates
- LongStream: Long-Sequence Streaming Autoregressive Visual Geometry
- Keyframe-Based Feed-Forward Visual Odometry
- MoE3D: A Mixture-of-Experts Module for 3D Reconstruction
- MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion
- VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction
- SLAMFormer-∞: Infinite SLAM Transformer for Unbounded Frontend and Backend Processing
- Dynamic Camera Poses and Where to Find Them
- PolyLayout: Multi-room Manhattan Layout Estimation
- The Fourth Monocular Depth Estimation Challenge
- LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning
- A Guide to Structureless Visual Localization
- UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
- Kitchen Robotic Manipulation utilizing Foundation Models
- MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors
- Label-Free Target-Domain Adaptation for Unconstrained Event-Image Feature Matching via Dual-Stage Distillation
- Any3DAvatar: Fast and High-Quality Full-Head 3D Avatar Reconstruction from Single Portrait Image
- SEAR: Simple and Efficient Adaptation of Visual Geometric Transformers for Unpaired RGB+Thermal 3D Reconstruction
- SmallGS: Gaussian Splatting-based Camera Pose Estimation for Small-Baseline Videos
- VSLAM-LAB: A Comprehensive Framework for Visual SLAM Methods and Datasets
- Pose Optimization for Autonomous Driving Datasets using Neural Rendering Models
- Towards Understanding Camera Motions in Any Video
- TAPIP3D: Tracking Any Point in Persistent 3D Geometry
- Back on Track: Bundle Adjustment for Dynamic Scene Reconstruction
- Mono3R: Exploiting Monocular Cues for Geometric 3D Reconstruction
- St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World
- Digital Twin Generation from Visual Data: A Survey
- AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis
- 3R-GS: Best Practice in Optimizing Camera Poses Along with 3DGS
- Regist3R: Incremental Registration with Stereo Foundation Model
- LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models
- Confidence matters: Leveraging Multi-view Geometric Priors for GS-based Reconstruction
- CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images
- Topometric Autonomous Vehicle Localization by Combining Visual Embeddings and Feed-Forward 3D Models
- Are Pretrained Image Matchers Good Enough for SAR-Optical Satellite Registration?
- Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting
- LL-Gaussian: Low-Light Scene Reconstruction and Enhancement via Gaussian Splatting for Novel View Synthesis
- GaussVideoDreamer: 3D Scene Generation with Video Diffusion and Inconsistency-Aware Gaussian Splatting
- Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
- BoxDreamer: Dreaming Box Corners for Generalizable Object Pose Estimation
- POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction
- D2USt3R: Enhancing 3D Reconstruction for Dynamic Scenes
- To Match or Not to Match: Revisiting Image Matching for Reliable Visual Place Recognition
- PanoDreamer: Consistent Text to 360-Degree Scene Generation
Related