vix.ing · top · new · best · stats · spec

Kentaro Tachibana

  1. PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-to-Speech Using Natural Language Descriptions
    2023/09/15 by Reo Shimizu, Shimizu, Reo, Ryuichi Yamamoto +11 · 18 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
  2. LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
    2024/06/12 by Masaya Kawamura, Ryuichi Yamamoto, Kawamura, Masaya +7 · 16 citations
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Subtitles and Audiovisual Media #Translation Studies and Practices #electronic engineering #information engineering
  3. Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform
    2022/10/28 by Masaya Kawamura, Kawamura, Masaya, Yuma Shirahata +5 · 6 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
  4. Phrase break prediction with bidirectional encoder representations in Japanese text-to-speech synthesis
    2021/04/26 by Kosuke Futamata, Futamata, Kosuke, Byeongseon Park +5 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  5. STUDIES: Corpus of Japanese Empathetic Dialogue Speech Towards Friendly Voice Agent
    2022/03/28 by Yuki Saito, Saito, Yuki, Yuto Nishimura +7 · 1 citation
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Natural Language Processing Techniques #Social Robot Interaction and HRI #Sound (cs.SD) #Speech and dialogue systems
  6. Audio-conditioned phonemic and prosodic annotation for building text-to-speech models from unlabeled speech data
    2024/06/12 by Yuma Shirahata, Shirahata, Yuma, Byeongseon Park +5 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  7. Cross-Speaker Emotion Transfer for Low-Resource Text-to-Speech Using Non-Parallel Voice Conversion with Pitch-Shift Data Augmentation
    2022/04/21 by Ryo Terashima, Terashima, Ryo, Ryuichi Yamamoto +11 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing
  8. CALLS: Japanese Empathetic Dialogue Speech Corpus of Complaint Handling and Attentive Listening in Customer Center
    2023/05/23 by Yuki Saito, Saito, Yuki, Eiji Iimori +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  9. DRSpeech: Degradation-Robust Text-to-Speech Synthesis with Frame-Level and Utterance-Level Acoustic Representation Learning
    2022/03/29 by Takaaki Saeki, Saeki, Takaaki, Kentaro Tachibana +3 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  10. Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
    2024/06/11 by T. IGARASHI, Yuki Saito, Igarashi, Takuto +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  11. Period VITS: Variational Inference with Explicit Pitch Modeling for End-to-end Emotional Speech Synthesis
    2022/10/28 by Yuma Shirahata, Ryuichi Yamamoto, Shirahata, Yuma +9 · 1 citation
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering