2025/01/18 by Jaekwon Im, Im, Jaekwon, Juhan Nam +1 · 3 citations
Computer Science · #Advanced Image Processing Techniques #Audio and Speech Processing (eess.AS) #Digital Filter Design and Implementation #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Signal Denoising Methods #Sound (cs.SD) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2501.10807
openalex publication_date 2025/01/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Versatile audio super-resolution (SR) is the challenging task of restoring high-frequency components from low-resolution audio with sampling rates between 4kHz and 32kHz in various domains such as music, speech, and sound effects. Previous diffusion-based SR methods suffer from slow inference due to the need for a large number of sampling steps. In this paper, we introduce FlashSR, a single-step diffusion model for versatile audio super-resolution aimed at producing 48kHz audio. FlashSR achieves fast inference by utilizing diffusion distillation with three objectives: distillation loss, adversarial loss, and distribution-matching distillation loss. We further enhance performance by proposing the SR Vocoder, which is specifically designed for SR models operating on mel-spectrograms. FlashSR demonstrates competitive performance with the current state-of-the-art model in both objective and subjective evaluations while being approximately 22 times faster.