2025/05/24 by Xiaobin Rong, Da‐Han Wang, Rong, Xiaobin +9 · 2 citations
Computer Science · Health Professions · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Background noise #Clipping (morphology) #Codec #Distortion (music) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Packet loss #Residual #Speech Recognition and Synthesis #Speech and Audio Processing #Speech coding #Speech enhancement #Voice activity detection #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2505.18533
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/05/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Universal speech enhancement aims to handle input speech with different distortions and input formats. To tackle this challenge, we present TS-URGENet, a Three-Stage Universal, Robust, and Generalizable speech Enhancement Network. To address various distortions, the proposed system employs a novel three-stage architecture consisting of a filling stage, a separation stage, and a restoration stage. The filling stage mitigates packet loss by preliminarily filling lost regions under noise interference, ensuring signal continuity. The separation stage suppresses noise, reverberation, and clipping distortion to improve speech clarity. Finally, the restoration stage compensates for bandwidth limitation, codec artifacts, and residual packet loss distortion, refining the overall speech quality. Our proposed TS-URGENet achieved outstanding performance in the Interspeech 2025 URGENT Challenge, ranking 2nd in Track 1.