Yuma Shirahata
- PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-to-Speech Using Natural Language Descriptions
2023/09/15 by Reo Shimizu, Ryuichi Yamamoto, Shimizu, Reo +11 · 18 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
- Universal Score-based Speech Enhancement with High Content Preservation
2024/06/18 by Robin Scheibler, Scheibler, Robin, Yusuke Fujita +5 · 19 citations
Computer Science · Engineering · #Speech and Audio Processing #Advanced Data Compression Techniques #Advanced Adaptive Filtering Techniques
- LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
2024/06/12 by Masaya Kawamura, Ryuichi Yamamoto, Kawamura, Masaya +7 · 16 citations
Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Subtitles and Audiovisual Media #Translation Studies and Practices #electronic engineering #information engineering
- Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform
2022/10/28 by Masaya Kawamura, Yuma Shirahata, Kawamura, Masaya +5 · 6 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
- Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
2024/06/12 by Yuma Shirahata, Shirahata, Yuma, Byeongseon Park +5 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Cross-Speaker Emotion Transfer for Low-Resource Text-to-Speech Using Non-Parallel Voice Conversion with Pitch-Shift Data Augmentation
2022/04/21 by Ryo Terashima, Ryuichi Yamamoto, Terashima, Ryo +11 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing
- Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
2025/06/05 by Hien Ohnaka, Ohnaka, Hien, Yuma Shirahata +5 · 2 citations
Computer Science · Psychology · #Speech Recognition and Synthesis #Emotion and Mood Recognition #Phonetics and Phonology Research
- Period VITS: Variational Inference with Explicit Pitch Modeling for End-to-end Emotional Speech Synthesis
2022/10/28 by Yuma Shirahata, Ryuichi Yamamoto, Shirahata, Yuma +9 · 1 citation
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering