2020/08/31 by Akash Gupta, Gupta, Akash, Abhishek Aich +3
Computer Science · Engineering · #Advanced Image Processing Techniques #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image Processing Techniques and Applications #Image and Video Processing (eess.IV) #Machine Learning (cs.LG) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2009.01005
openalex publication_date 2020/08/31 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
Existing works address the problem of generating high frame-rate sharp videos\nby separately learning the frame deblurring and frame interpolation modules.\nMost of these approaches have a strong prior assumption that all the input\nframes are blurry whereas in a real-world setting, the quality of frames\nvaries. Moreover, such approaches are trained to perform either of the two\ntasks - deblurring or interpolation - in isolation, while many practical\nsituations call for both. Different from these works, we address a more\nrealistic problem of high frame-rate sharp video synthesis with no prior\nassumption that input is always blurry. We introduce a novel architecture,\nAdaptive Latent Attention Network (ALANET), which synthesizes sharp high\nframe-rate videos with no prior knowledge of input frames being blurry or not,\nthereby performing the task of both deblurring and interpolation. We\nhypothesize that information from the latent representation of the consecutive\nframes can be utilized to generate optimized representations for both frame\ndeblurring and frame interpolation. Specifically, we employ combination of\nself-attention and cross-attention module between consecutive frames in the\nlatent space to generate optimized representation for each frame. The optimized\nrepresentation learnt using these attention modules help the model to generate\nand interpolate sharp frames. Extensive experiments on standard datasets\ndemonstrate that our method performs favorably against various state-of-the-art\napproaches, even though we tackle a much more difficult problem.\n