vix.ing · top · new · best · stats · spec

Coherent 3D Scene Diffusion From a Single RGB Image

2024/12/13 by Manuel Dahnert, Dahnert, Manuel, Angela Dai +5 · 3 citations
Computer Science · #Advanced Vision and Imaging #Computer Graphics and Visualization Techniques #Optical measurement and interference techniques

paper · pdf · doi:10.48550/arxiv.2412.10294

Abstract

We present a novel diffusion-based approach for coherent 3D scene\nreconstruction from a single RGB image. Our method utilizes an\nimage-conditioned 3D scene diffusion model to simultaneously denoise the 3D\nposes and geometries of all objects within the scene. Motivated by the\nill-posed nature of the task and to obtain consistent scene reconstruction\nresults, we learn a generative scene prior by conditioning on all scene objects\nsimultaneously to capture the scene context and by allowing the model to learn\ninter-object relationships throughout the diffusion process. We further propose\nan efficient surface alignment loss to facilitate training even in the absence\nof full ground-truth annotation, which is common in publicly available\ndatasets. This loss leverages an expressive shape representation, which enables\ndirect point sampling from intermediate shape predictions. By framing the task\nof single RGB image 3D scene reconstruction as a conditional diffusion process,\nour approach surpasses current state-of-the-art methods, achieving a 12.04%\nimprovement in AP3D on SUN RGB-D and a 13.43% increase in F-Score on Pix3D.\n

Cited by

Related