vix.ing · top · new · best · stats · spec

CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis

2025/09/08 by Daniel Watson, Kong, Xin, Yannick Strümpler +6 · 2 citations
Computer Science · Engineering · #Advanced Optical Imaging Technologies #Advanced Vision and Imaging #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences

paper · pdf · doi:10.48550/arxiv.2509.06579

openalex publication_date 2025/09/08 · openalex created_date 2025/10/11 · openalex updated_date 2026/07/28

Abstract

Multi-view diffusion models have shown promise in 3D novel view synthesis, but most existing methods adopt a non-autoregressive formulation. This limits their applicability in world modeling, as they only support a fixed number of views and suffer from slow inference due to denoising all frames simultaneously. To address these limitations, we propose CausNVS, a multi-view diffusion model in an autoregressive setting, which supports arbitrary input-output view configurations and generates views sequentially. We train CausNVS with causal masking and per-frame noise, using pairwise-relative camera pose encodings (CaPE) for precise camera control. At inference time, we combine a spatially-aware sliding-window with key-value caching and noise conditioning augmentation to mitigate drift. Our experiments demonstrate that CausNVS supports a broad range of camera trajectories, enables flexible autoregressive novel view synthesis, and achieves consistently strong visual quality across diverse settings. Project page: https://kxhit.github.io/CausNVS.html.

Citations

Cited by

Related