2025/08/04 by Yulun Wu, Wu, Yulun, Zhongweiyang Xu +6
Computer Science · Neuroscience · #Anechoic chamber #Constraint (computer-aided design) #Diffusion #Hearing Loss and Rehabilitation #Impulse (physics) #Impulse response #Microphone #Reverberation #Sampling (signal processing) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech enhancement
paper · open access · doi:10.48550/arxiv.2508.02071
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/08/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We consider the problem of multi-channel single-speaker blind dereverberation, where multi-channel mixtures are used to recover the clean anechoic speech. To solve this problem, we propose USD-DPS, Unsupervised Speech Dereverberation via Diffusion Posterior Sampling. USD-DPS uses an unconditional clean speech diffusion model as a strong prior to solve the problem by posterior sampling. At each diffusion sampling step, we estimate all microphone channels' room impulse responses (RIRs), which are further used to enforce a multi-channel mixture consistency constraint for diffusion guidance. For multi-channel RIR estimation, we estimate reference-channel RIR by optimizing RIR parameters of a sub-band RIR signal model, with the Adam optimizer. We estimate non-reference channels' RIRs analytically using forward convolutive prediction (FCP). We found that this combination provides a good balance between sampling efficiency and RIR prior modeling, which shows superior performance among unsupervised dereverberation approaches. An audio demo page is provided in https://usddps.github.io/USDDPSdemo/.