vix.ing · top · new · best · stats · spec

Personalized PercepNet: Real-time, Low-complexity Target Voice\n Separation and Enhancement

2021/06/08 by Ritwik Giri, Shrikant Venkataramani, Giri, Ritwik +7
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2106.04129

openalex publication_date 2021/06/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The presence of multiple talkers in the surrounding environment poses a\ndifficult challenge for real-time speech communication systems considering the\nconstraints on network size and complexity. In this paper, we present\nPersonalized PercepNet, a real-time speech enhancement model that separates a\ntarget speaker from a noisy multi-talker mixture without compromising on\ncomplexity of the recently proposed PercepNet. To enable speaker-dependent\nspeech enhancement, we first show how we can train a perceptually motivated\nspeaker embedder network to produce a representative embedding vector for the\ngiven speaker. Personalized PercepNet uses the target speaker embedding as\nadditional information to pick out and enhance only the target speaker while\nsuppressing all other competing sounds. Our experiments show that the proposed\nmodel significantly outperforms PercepNet and other baselines, both in terms of\nobjective speech enhancement metrics and human opinion scores.\n

Related