vix.ing · top · new · best · stats · spec

Yuma Shirahata

  1. PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-to-Speech Using Natural Language Descriptions
    2023/09/15 by Reo Shimizu, Ryuichi Yamamoto, Shimizu, Reo +11 · 18 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
  2. Universal Score-based Speech Enhancement with High Content Preservation
    2024/06/18 by Robin Scheibler, Scheibler, Robin, Yusuke Fujita +5 · 19 citations
    Computer Science · Engineering · #Speech and Audio Processing #Advanced Data Compression Techniques #Advanced Adaptive Filtering Techniques
  3. LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
    2024/06/12 by Masaya Kawamura, Ryuichi Yamamoto, Kawamura, Masaya +7 · 16 citations
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Subtitles and Audiovisual Media #Translation Studies and Practices #electronic engineering #information engineering
  4. Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform
    2022/10/28 by Masaya Kawamura, Yuma Shirahata, Kawamura, Masaya +5 · 6 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
  5. Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
    2024/06/12 by Yuma Shirahata, Shirahata, Yuma, Byeongseon Park +5 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  6. Cross-Speaker Emotion Transfer for Low-Resource Text-to-Speech Using Non-Parallel Voice Conversion with Pitch-Shift Data Augmentation
    2022/04/21 by Ryo Terashima, Ryuichi Yamamoto, Terashima, Ryo +11 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing
  7. Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
    2025/06/05 by Hien Ohnaka, Ohnaka, Hien, Yuma Shirahata +5 · 2 citations
    Computer Science · Psychology · #Speech Recognition and Synthesis #Emotion and Mood Recognition #Phonetics and Phonology Research
  8. Period VITS: Variational Inference with Explicit Pitch Modeling for End-to-end Emotional Speech Synthesis
    2022/10/28 by Yuma Shirahata, Ryuichi Yamamoto, Shirahata, Yuma +9 · 1 citation
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering