2021/10/29 by Alejandro Pérez-González-de-Martos, Pérez-González-de-Martos, Alejandro, Albert Sanchís +4 · 2 citations
Computer Science · Engineering · Mathematics · #Artificial intelligence #Audio and Speech Processing (eess.AS) #Autoregressive model #Computer science #Engineering #FOS: Computer and information sciences #FOS: Electrical engineering #Hidden Markov model #Mathematics #Natural Language Processing Techniques #Natural language processing #Naturalness #Pipeline (software) #Programming language #Quality (philosophy) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech recognition #Statistics #Task (project management) #Topic Modeling #cs.SD #eess.AS #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2110.15792
published in arXiv (Cornell University) (Cornell University)
arxiv created 2021/10/29 · openalex publication_date 2021/10/29 · arxiv updated 2021/11/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
This paper presents the VRAIN-UPV MLLP's speech synthesis system for the SH1 task of the Blizzard Challenge 2021. The SH1 task consisted in building a Spanish text-to-speech system trained on (but not limited to) the corpus released by the Blizzard Challenge 2021 organization. It included 5 hours of studio-quality recordings from a native Spanish female speaker. In our case, this dataset was solely used to build a two-stage neural text-to-speech pipeline composed of a non-autoregressive acoustic model with explicit duration modeling and a HiFi-GAN neural vocoder. Our team is identified as J in the evaluation results. Our system obtained very good results in the subjective evaluation tests. Only one system among other 11 participants achieved better naturalness than ours. Concretely, it achieved a naturalness MOS of 3.61 compared to 4.21 for real samples.