2021/03/16 by Siyang Yuan, Pengyu Cheng, Yuan, Siyang +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2103.09420
openalex publication_date 2021/03/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Voice style transfer, also called voice conversion, seeks to modify one\nspeaker's voice to generate speech as if it came from another (target) speaker.\nPrevious works have made progress on voice conversion with parallel training\ndata and pre-known speakers. However, zero-shot voice style transfer, which\nlearns from non-parallel data and generates voices for previously unseen\nspeakers, remains a challenging problem. We propose a novel zero-shot voice\ntransfer method via disentangled representation learning. The proposed method\nfirst encodes speaker-related style and voice content of each input voice into\nseparated low-dimensional embedding spaces, and then transfers to a new voice\nby combining the source content embedding and target style embedding through a\ndecoder. With information-theoretic guidance, the style and content embedding\nspaces are representative and (ideally) independent of each other. On\nreal-world VCTK datasets, our method outperforms other baselines and obtains\nstate-of-the-art results in terms of transfer accuracy and voice naturalness\nfor voice style transfer experiments under both many-to-many and zero-shot\nsetups.\n