vix.ing · top · new · best · stats

TS-URGENet: A Three-stage Universal Robust and Generalizable Speech Enhancement Network

2025/05/24 by Xiaobin Rong, Da‐Han Wang, Rong, Xiaobin +9 · 2 citations
Computer Science · Health Professions · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Background noise #Clipping (morphology) #Codec #Distortion (music) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Packet loss #Residual #Speech Recognition and Synthesis #Speech and Audio Processing #Speech coding #Speech enhancement #Voice activity detection #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2505.18533

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2025/05/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Universal speech enhancement aims to handle input speech with different distortions and input formats. To tackle this challenge, we present TS-URGENet, a Three-Stage Universal, Robust, and Generalizable speech Enhancement Network. To address various distortions, the proposed system employs a novel three-stage architecture consisting of a filling stage, a separation stage, and a restoration stage. The filling stage mitigates packet loss by preliminarily filling lost regions under noise interference, ensuring signal continuity. The separation stage suppresses noise, reverberation, and clipping distortion to improve speech clarity. Finally, the restoration stage compensates for bandwidth limitation, codec artifacts, and residual packet loss distortion, refining the overall speech quality. Our proposed TS-URGENet achieved outstanding performance in the Interspeech 2025 URGENT Challenge, ranking 2nd in Track 1.

Citations

Cited by

Related