GTR: Gaussian Splatting Tracking and Reconstruction of Unknown Objects Based on Appearance and Geometric Complexity
2025/05/17 by Takuya Ikeda, Ikeda, Takuya, Sergey Zakharov +19
Computer Science · Engineering · #3D Shape Modeling and Analysis #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Robotics (cs.RO) #Video Surveillance and Tracking Methods
paper · pdf · doi:10.48550/arxiv.2505.11905
openalex publication_date 2025/05/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We present a novel method for 6-DoF object tracking and high-quality 3D reconstruction from monocular RGBD video. Existing methods, while achieving impressive results, often struggle with complex objects, particularly those exhibiting symmetry, intricate geometry or complex appearance. To bridge these gaps, we introduce an adaptive method that combines 3D Gaussian Splatting, hybrid geometry/appearance tracking, and key frame selection to achieve robust tracking and accurate reconstructions across a diverse range of objects. Additionally, we present a benchmark covering these challenging object classes, providing high-quality annotations for evaluating both tracking and reconstruction performance. Our approach demonstrates strong capabilities in recovering high-fidelity object meshes, setting a new standard for single-sensor 3D reconstruction in open-world environments.
Citations
- SAM 2: Segment Anything in Images and Videos
- FlowMap: High-Quality Camera Poses, Intrinsics, and Depth via Gradient Descent
- Zero-Shot Multi-Object Scene Completion
- Cameras as Rays: Pose Estimation via Ray Diffusion
- DiffusionNOCS: Managing Symmetry and Uncertainty in Sim2Real Multi-Modal Category-level Pose Estimation
- BootsTAP: Bootstrapped Training for Tracking-Any-Point
- ZeroShape: Regression-based Zero-shot Shape Reconstruction
- FoundationPose-TensorRT
- Gaussian Splatting SLAM
- Gaussian-SLAM: Photo-realistic Dense SLAM with Gaussian Splatting
- SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM
- FreeZe: Training-free zero-shot 6D pose estimation with geometric and vision foundation models
- FoundPose: Unseen Object Pose Estimation with Foundation Features
- FSD: Fast Self-Supervised Single RGB-D to Categorical 3D Objects
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- GenPose: Generative Category-level Object Pose Estimation via Diffusion Models
- BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects
- Fusing Visual Appearance and Geometry for Multi-modality 6DoF Object Tracking
- vMAP: Vectorised Object Mapping for Neural Field SLAM
- PyPose: A Library for Robot Learning with Physics-based Optimization
- Object-Compositional Neural Implicit Surfaces
- CenterSnap: Single-Shot Multi-Object 3D Shape Reconstruction and Categorical 6D Pose and Size Estimation
- NICE-SLAM: Neural Implicit Scalable Encoding for SLAM
- Extracting Triangular 3D Models, Materials, and Lighting From Images
- A Learned Stereo Depth System for Robotic Manipulation in Homes
- Learning Object-Compositional Neural Radiance Field for Editable Scene Rendering
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras
- BundleTrack: 6D Pose Tracking for Novel Objects without Instance or Category-Level 3D Models
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction
- UNISURF: Unifying Neural Implicit Surfaces and Radiance Fields for Multi-View Reconstruction
- LoFTR: Detector-Free Local Feature Matching with Transformers
- Pix2Surf: Learning Parametric 3D Surface Models of Objects from Images
- Shape Prior Deformation for Categorical 6D Object Pose and Size Estimation
- MeshSDF: Differentiable Iso-Surface Extraction
- EPOS: Estimating 6D Pose of Objects with Symmetries
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- Dynamic SLAM: The Need For Speed
- Learning Canonical Shape Space for Category-Level 6D Object Pose and Size Estimation
- DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion
- Implicit 3D Orientation Learning for 6D Object Detection from RGB Images
- Real-Time Seamless Single Shot 6D Object Pose Prediction
- PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes
- SSD-6D: Making RGB-based 3D detection and 6D pose estimation great again
- BB8: A Scalable, Accurate, Robust to Partial Occlusion Method for Predicting the 3D Poses of Challenging Objects without Using Depth
- Deep Learning of Local RGB-D Patches for 3D Object Detection and 6D Pose Estimation
- Image quality assessment: from error visibility to structural similarity
- ShAPO: Implicit Representations for Multi-Object Shape, Appearance, and Pose Optimization
Related