vix.ing · top · new · best · stats · spec

DC-VSR: Spatially and Temporally Consistent Video Super-Resolution with Video Diffusion Prior

2025/02/05 by Gyujin Sim, Han, Janghyeok, Sim, Gyujin +10 · 1 citation
Computer Science · #Advanced Image Processing Techniques #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #FOS: Electrical engineering #Graphics (cs.GR) #Image and Signal Denoising Methods #Image and Video Processing (eess.IV) #Image and Video Quality Assessment #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2502.03502

openalex publication_date 2025/02/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Video super-resolution (VSR) aims to reconstruct a high-resolution (HR) video from a low-resolution (LR) counterpart. Achieving successful VSR requires producing realistic HR details and ensuring both spatial and temporal consistency. To restore realistic details, diffusion-based VSR approaches have recently been proposed. However, the inherent randomness of diffusion, combined with their tile-based approach, often leads to spatio-temporal inconsistencies. In this paper, we propose DC-VSR, a novel VSR approach to produce spatially and temporally consistent VSR results with realistic textures. To achieve spatial and temporal consistency, DC-VSR adopts a novel Spatial Attention Propagation (SAP) scheme and a Temporal Attention Propagation (TAP) scheme that propagate information across spatio-temporal tiles based on the self-attention mechanism. To enhance high-frequency details, we also introduce Detail-Suppression Self-Attention Guidance (DSSAG), a novel diffusion guidance scheme. Comprehensive experiments demonstrate that DC-VSR achieves spatially and temporally consistent, high-quality VSR results, outperforming previous approaches.

Cited by

Related