vix.ing · top · new · best · stats

Video Inpainting by Jointly Learning Temporal Structure and Spatial Details

2018/06/22 by Chuan Wang, Wang, Chuan, Haibin Huang +5 · 12 citations
Computer Science · #Advanced Image Processing Techniques #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #cs.CV

paper · pdf · doi:10.48550/arxiv.1806.08482

Accepted by AAAI 2019

openalex publication_date 2018/06/22 · arxiv created 2018/12/03 · arxiv updated 2018/12/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We present a new data-driven video inpainting method for recovering missing regions of video frames. A novel deep learning architecture is proposed which contains two sub-networks: a temporal structure inference network and a spatial detail recovering network. The temporal structure inference network is built upon a 3D fully convolutional architecture: it only learns to complete a low-resolution video volume given the expensive computational cost of 3D convolution. The low resolution result provides temporal guidance to the spatial detail recovering network, which performs image-based inpainting with a 2D fully convolutional network to produce recovered video frames in their original resolution. Such two-step network design ensures both the spatial quality of each frame and the temporal coherence across frames. Our method jointly trains both sub-networks in an end-to-end manner. We provide qualitative and quantitative evaluation on three datasets, demonstrating that our method outperforms previous learning-based video inpainting methods.

Citations

Cited by

Related