vix.ing · top · new · best · stats · spec

Takamichi, Shinnosuke

  1. UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
    2022/04/05 by Takaaki Saeki, Saeki, Takaaki, Detai Xin +9 · 192 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
  2. JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis
    2017/10/28 by Ryosuke Sonobe, Sonobe, Ryosuke, Shinnosuke Takamichi +3 · 23 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  3. SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
    2024/01/30 by Takaaki Saeki, Saeki, Takaaki, Soumi Maiti +7 · 33 citations
    Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  4. BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
    2024/09/09 by Detai Xin, Xin, Detai, Xu Tan +5 · 35 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  5. YODAS: Youtube-Oriented Dataset for Audio and Speech
    2024/06/02 by Xinjian Li, Li, Xinjian, Shinnosuke Takamichi +9 · 30 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Video Analysis and Summarization #electronic engineering #information engineering
  6. JVS corpus: free Japanese multi-speaker voice corpus
    2019/08/17 by Shinnosuke Takamichi, Kentaro Mitsui, Takamichi, Shinnosuke +9 · 12 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  7. SpoofCeleb: Speech Deepfake Detection and SASV In The Wild
    2024/09/18 by Jee-weon Jung, Jung, Jee-weon, Yihan Wu +25 · 24 citations
    Computer Science · #Speech Recognition and Synthesis
  8. PJS: phoneme-balanced Japanese singing voice corpus
    2020/06/04 by Junya Koguchi, Shinnosuke Takamichi, Koguchi, Junya +1 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. JVS-MuSiC: Japanese multispeaker singing-voice corpus
    2020/01/20 by Hiroki Tamaru, Tamaru, Hiroki, Shinnosuke Takamichi +5 · 7 citations
    Computer Science · Medicine · #Speech Recognition and Synthesis #Music and Audio Processing #Voice and Speech Disorders
  10. Phase reconstruction from amplitude spectrograms based on von-Mises-distribution deep neural network
    2018/07/10 by Shinnosuke Takamichi, Yuki Saito, Takamichi, Shinnosuke +7 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  11. JTubeSpeech: corpus of Japanese speech collected from YouTube for speech recognition and speaker verification
    2021/12/17 by Takamichi, Shinnosuke, Kürzinger, Ludwig, Saeki, Takaaki +2 · 6 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  12. ESPnet2-TTS: Extending the Edge of TTS Research
    2021/10/15 by Tomoki Hayashi, Ryuichi Yamamoto, Hayashi, Tomoki +17 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  13. RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
    2024/04/04 by Detai Xin, Xu Tan, Xin, Detai +19 · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  14. Voice Conversion Using Sequence-to-Sequence Learning of Context Posterior Probabilities
    2017/04/10 by Hiroyuki Miyoshi, Miyoshi, Hiroyuki, Yuki Saito +5 · 4 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Topic Modeling
  15. Coco-Nut: Corpus of Japanese Utterance and Voice Characteristics Description for Prompt-based Control
    2023/09/24 by Aya Watanabe, Watanabe, Aya, Shinnosuke Takamichi +9 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  16. Text-To-Speech Synthesis In The Wild
    2024/09/13 by Jee-weon Jung, Wangyou Zhang, Jung, Jee-weon +25 · 6 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering
  17. JNV Corpus: A Corpus of Japanese Nonverbal Vocalizations with Diverse Phrases and Emotions
    2023/05/21 by Detai Xin, Shinnosuke Takamichi, Xin, Detai +3 · 3 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  18. Learning to Speak from Text: Zero-Shot Multilingual Text-to-Speech with Unsupervised Text Pretraining
    2023/01/30 by Takaaki Saeki, Soumi Maiti, Saeki, Takaaki +9 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  19. J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis
    2022/01/26 by Shinnosuke Takamichi, Wataru Nakata, Takamichi, Shinnosuke +5 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  20. SelfRemaster: Self-Supervised Speech Restoration with Analysis-by-Synthesis Approach Using Channel Modeling
    2022/03/24 by Takaaki Saeki, Saeki, Takaaki, Shinnosuke Takamichi +7 · 2 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis
  21. Text-to-speech synthesis from dark data with evaluation-in-the-loop data selection
    2022/10/26 by Kentaro Seki, Shinnosuke Takamichi, Seki, Kentaro +5 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  22. Improving Speech Prosody of Audiobook Text-to-Speech Synthesis with Acoustic and Textual Contexts
    2022/11/04 by Xin, Detai, Adavanne, Sharath, Ang, Federico +3 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  23. Statistical Parametric Speech Synthesis Incorporating Generative Adversarial Networks
    2017/09/23 by Yuki Saito, Saito, Yuki, Shinnosuke Takamichi +3 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  24. RWCP-SSD-Onomatopoeia: Onomatopoeic Word Dataset for Environmental Sound Synthesis
    2020/07/09 by Yuki Okamoto, Okamoto, Yuki, Keisuke Imoto +9 · 1 citation
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Animal Vocal Communication and Behavior #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  25. JSSS: free Japanese speech corpus for summarization and simplification
    2020/10/05 by Shinnosuke Takamichi, Mamoru Komachi, Takamichi, Shinnosuke +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Text Readability and Simplification #Topic Modeling #electronic engineering #information engineering
  26. STUDIES: Corpus of Japanese Empathetic Dialogue Speech Towards Friendly Voice Agent
    2022/03/28 by Yuki Saito, Saito, Yuki, Yuto Nishimura +7 · 1 citation
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Natural Language Processing Techniques #Social Robot Interaction and HRI #Sound (cs.SD) #Speech and dialogue systems
  27. Do learned speech symbols follow Zipf's law?
    2023/09/18 by Shinnosuke Takamichi, Hiroki Maeda, Takamichi, Shinnosuke +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  28. CALLS: Japanese Empathetic Dialogue Speech Corpus of Complaint Handling and Attentive Listening in Customer Center
    2023/05/23 by Saito, Yuki, Iimori, Eiji, Takamichi, Shinnosuke +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  29. Improving robustness of spontaneous speech synthesis with linguistic speech regularization and pseudo-filled-pause insertion
    2022/10/18 by Yuta Matsunaga, Takaaki Saeki, Matsunaga, Yuta +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  30. Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue History
    2022/06/16 by Nishimura, Yuto, Saito, Yuki, Takamichi, Shinnosuke +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  31. Mid-attribute speaker generation using optimal-transport-based interpolation of Gaussian mixture models
    2022/10/18 by Aya Watanabe, Shinnosuke Takamichi, Watanabe, Aya +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  32. Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus
    2023/05/21 by Xin, Detai, Takamichi, Shinnosuke, Morimatsu, Ai +1 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  33. Diversity-based core-set selection for text-to-speech with linguistic and acoustic features
    2023/09/15 by Kentaro Seki, Seki, Kentaro, Shinnosuke Takamichi +5 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  34. J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
    2024/07/22 by Nakata, Wataru, Seki, Kentaro, Yanaka, Hitomi +3 · 3 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  35. V2S attack: building DNN-based voice conversion from automatic speaker verification
    2019/08/05 by Taiki Nakamura, Nakamura, Taiki, Yuki Saito +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  36. Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
    2024/07/05 by Hitoshi Suda, Aya Watanabe, Suda, Hitoshi +3 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  37. SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
    2024/06/11 by Saito, Yuki, Igarashi, Takuto, Seki, Kentaro +4 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  38. How Should We Evaluate Synthesized Environmental Sounds
    2022/08/16 by Yuki Okamoto, Keisuke Imoto, Okamoto, Yuki +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Sound (cs.SD) #electronic engineering #information engineering
  39. Building speech corpus with diverse voice characteristics for its prompt-based representation
    2024/03/20 by Aya Watanabe, Shinnosuke Takamichi, Watanabe, Aya +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  40. Environmental sound synthesis from vocal imitations and sound event labels
    2023/04/29 by Okamoto, Yuki, Imoto, Keisuke, Takamichi, Shinnosuke +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  41. Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
    2024/06/11 by T. IGARASHI, Igarashi, Takuto, Yuki Saito +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  42. TTSOps: A Closed-Loop Corpus Optimization Framework for Training Multi-Speaker TTS Models from Dark Data
    2025/06/18 by Kentaro Seki, Seki, Kentaro, Shinnosuke Takamichi +5 · 4 citations
    Computer Science · #FOS: Computer and information sciences #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling
  43. DNN-based Speaker Embedding Using Subjective Inter-speaker Similarity for Multi-speaker Modeling in Speech Synthesis
    2019/07/19 by Yuki Saito, Saito, Yuki, Shinnosuke Takamichi +3 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing