Kentaro Tachibana
- PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-to-Speech Using Natural Language Descriptions
2023/09/15 by Reo Shimizu, Shimizu, Reo, Ryuichi Yamamoto +11 · 18 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
- LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
2024/06/12 by Masaya Kawamura, Ryuichi Yamamoto, Kawamura, Masaya +7 · 16 citations
Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Subtitles and Audiovisual Media #Translation Studies and Practices #electronic engineering #information engineering
- Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform
2022/10/28 by Masaya Kawamura, Kawamura, Masaya, Yuma Shirahata +5 · 6 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
- Phrase break prediction with bidirectional encoder representations in Japanese text-to-speech synthesis
2021/04/26 by Kosuke Futamata, Futamata, Kosuke, Byeongseon Park +5 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- STUDIES: Corpus of Japanese Empathetic Dialogue Speech Towards Friendly Voice Agent
2022/03/28 by Yuki Saito, Saito, Yuki, Yuto Nishimura +7 · 1 citation
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Natural Language Processing Techniques #Social Robot Interaction and HRI #Sound (cs.SD) #Speech and dialogue systems
- Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
2024/06/12 by Yuma Shirahata, Shirahata, Yuma, Byeongseon Park +5 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Cross-Speaker Emotion Transfer for Low-Resource Text-to-Speech Using Non-Parallel Voice Conversion with Pitch-Shift Data Augmentation
2022/04/21 by Ryo Terashima, Terashima, Ryo, Ryuichi Yamamoto +11 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing
- CALLS: Japanese Empathetic Dialogue Speech Corpus of Complaint Handling and Attentive Listening in Customer Center
2023/05/23 by Yuki Saito, Saito, Yuki, Eiji Iimori +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- DRSpeech: Degradation-Robust Text-to-Speech Synthesis with Frame-Level and Utterance-Level Acoustic Representation Learning
2022/03/29 by Takaaki Saeki, Saeki, Takaaki, Kentaro Tachibana +3 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
2024/06/11 by T. IGARASHI, Yuki Saito, Igarashi, Takuto +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Period VITS: Variational Inference with Explicit Pitch Modeling for End-to-end Emotional Speech Synthesis
2022/10/28 by Yuma Shirahata, Ryuichi Yamamoto, Shirahata, Yuma +9 · 1 citation
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering