LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
2025/12/15 by Ding, Tianye, Xie, Yiming, Liang, Yiqing +3
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2512.13680
Abstract
Recent feed-forward reconstruction models like VGGT and π3 achieve impressive reconstruction quality but cannot process streaming videos due to quadratic memory complexity, limiting their practical deployment. While existing streaming methods address this through learned memory mechanisms or causal attention, they require extensive retraining and may not fully leverage the strong geometric priors of state-of-the-art offline models. We propose LASER, a training-free framework that converts an offline reconstruction model into a streaming system by aligning predictions across consecutive temporal windows. We observe that simple similarity transformation (Sim(3)) alignment fails due to layer depth misalignment: monocular scale ambiguity causes relative depth scales of different scene layers to vary inconsistently between windows. To address this, we introduce layer-wise scale alignment, which segments depth predictions into discrete layers, computes per-layer scale factors, and propagates them across both adjacent windows and timestamps. Extensive experiments show that LASER achieves state-of-the-art performance on camera pose estimation and point map reconstruction %quality with offline models while operating at 14 FPS with 6 GB peak memory on a RTX A6000 GPU, enabling practical deployment for kilometer-scale streaming videos. Project website: \hrefhttps://neu-vi.github.io/LASER/https://neu-vi.github.io/LASER/
Citations
- TTT3R: 3D Reconstruction as Test-Time Training
- WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
- STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
- LONG3R: Long Sequence Streaming 3D Reconstruction
- VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
- Streaming 4D Visual Geometry Transformer
- Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory
- VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
- St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World
- D2USt3R: Enhancing 3D Reconstruction for Dynamic Scenes
- Easi3R: Estimating Disentangled Motion from DUSt3R Without Training
- VGGT: Visual Geometry Grounded Transformer
- Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
- Continuous 3D Perception Model with Persistent State
- Zero-Shot Monocular Scene Flow Estimation in the Wild
- MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors
- Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos
- Monocular Dynamic Gaussian Splatting: Fast, Brittle, and Scene Complexity Rules
- MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
- Align3R: Aligned Monocular Depth Estimation for Dynamic Videos
- MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
- 3D Reconstruction with Spatial Memory
- Deep Patch Visual SLAM
- SAM 2: Segment Anything in Images and Videos
- Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-aware Spatio-Temporal Sampling
- DUSt3R: Geometric 3D Vision Made Easy
- GauFRe: Gaussian Deformation Fields for Real-time Dynamic Novel View Synthesis
- 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
- Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction
- Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- K-Planes: Explicit Radiance Fields in Space, Time, and Appearance
- HexPlane: A Fast Representation for Dynamic Scenes
- Deep Patch Visual Odometry
- MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction
- Efficient Geometry-aware 3D Generative Adversarial Networks
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction
- NeuralRecon: Real-Time Coherent 3D Reconstruction from Monocular Video
- D-NeRF: Neural Radiance Fields for Dynamic Scenes
- Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic Scenes
- Nerfies: Deformable Neural Radiance Fields
- Atlas: End-to-End 3D Scene Reconstruction from Posed Images
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- LDSO: Direct Sparse Odometry with Loop Closure
- MVSNet: Depth Inference for Unstructured Multi-view Stereo
- DeepMVS: Learning Multi-view Stereopsis
- The 2017 DAVIS Challenge on Video Object Segmentation
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- π3: Permutation-Equivariant Visual Geometry Learning
- A Simplex Method for Function Minimization
Related