vix.ing · top · new · best · stats

Self-Supervised Equivariant Scene Synthesis from Video

2021/02/01 by Cinjon Resnick, Or Litany, Resnick, Cinjon +9
Computer Science · #Advanced Vision and Imaging #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #cs.CV

paper · pdf · doi:10.48550/arxiv.2102.00863

arXiv admin note: text overlap with arXiv:2011.05787

arxiv created 2021/02/01 · openalex publication_date 2021/02/01 · arxiv updated 2021/02/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We propose a self-supervised framework to learn scene representations from video that are automatically delineated into background, characters, and their animations. Our method capitalizes on moving characters being equivariant with respect to their transformation across frames and the background being constant with respect to that same transformation. After training, we can manipulate image encodings in real time to create unseen combinations of the delineated components. As far as we know, we are the first method to perform unsupervised extraction and synthesis of interpretable background, character, and animation. We demonstrate results on three datasets: Moving MNIST with backgrounds, 2D video game sprites, and Fashion Modeling.

Citations

Related