vix.ing · top · new · best · stats · spec

Adjusting Pleasure-Arousal-Dominance for Continuous Emotional\n Text-to-speech Synthesizer

2019/06/13 by Azam Rabiee, Taeho Kim, Rabiee, Azam +3 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Neural Networks and Applications #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1906.05507

openalex publication_date 2019/06/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Emotion is not limited to discrete categories of happy, sad, angry, fear,\ndisgust, surprise, and so on. Instead, each emotion category is projected into\na set of nearly independent dimensions, named pleasure (or valence), arousal,\nand dominance, known as PAD. The value of each dimension varies from -1 to 1,\nsuch that the neutral emotion is in the center with all-zero values. Training\nan emotional continuous text-to-speech (TTS) synthesizer on the independent\ndimensions provides the possibility of emotional speech synthesis with\nunlimited emotion categories. Our end-to-end neural speech synthesizer is based\non the well-known Tacotron. Empirically, we have found the optimum network\narchitecture for injecting the 3D PADs. Moreover, the PAD values are adjusted\nfor the speech synthesis purpose.\n

Citations

Cited by

Related