vix.ing · top · new · best · stats · spec

RUST: Latent Neural Scene Representations from Unposed Imagery

2022/11/25 by Mehdi S. M. Sajjadi, Sajjadi, Mehdi S. M., Aravindh Mahendran +11 · 6 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Graphics (cs.GR) #Image and Video Processing (eess.IV) #Machine Learning (cs.LG) #Robotics and Sensor-Based Localization #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2211.14306

openalex publication_date 2022/11/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Inferring the structure of 3D scenes from 2D observations is a fundamental challenge in computer vision. Recently popularized approaches based on neural scene representations have achieved tremendous impact and have been applied across a variety of applications. One of the major remaining challenges in this space is training a single model which can provide latent representations which effectively generalize beyond a single scene. Scene Representation Transformer (SRT) has shown promise in this direction, but scaling it to a larger set of diverse scenes is challenging and necessitates accurately posed ground truth data. To address this problem, we propose RUST (Really Unposed Scene representation Transformer), a pose-free approach to novel view synthesis trained on RGB images alone. Our main insight is that one can train a Pose Encoder that peeks at the target image and learns a latent pose embedding which is used by the decoder for view synthesis. We perform an empirical investigation into the learned latent pose structure and show that it allows meaningful test-time camera transformations and accurate explicit pose readouts. Perhaps surprisingly, RUST achieves similar quality as methods which have access to perfect camera pose, thereby unlocking the potential for large-scale training of amortized neural scene representations.

Cited by

Related