Takamichi, Shinnosuke
- UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
2022/04/05 by Takaaki Saeki, Saeki, Takaaki, Detai Xin +9 · 192 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
- JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis
2017/10/28 by Ryosuke Sonobe, Sonobe, Ryosuke, Shinnosuke Takamichi +3 · 23 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
2024/01/30 by Takaaki Saeki, Saeki, Takaaki, Soumi Maiti +7 · 33 citations
Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
- BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
2024/09/09 by Detai Xin, Xin, Detai, Xu Tan +5 · 35 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- YODAS: Youtube-Oriented Dataset for Audio and Speech
2024/06/02 by Xinjian Li, Li, Xinjian, Shinnosuke Takamichi +9 · 30 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Video Analysis and Summarization #electronic engineering #information engineering
- JVS corpus: free Japanese multi-speaker voice corpus
2019/08/17 by Shinnosuke Takamichi, Kentaro Mitsui, Takamichi, Shinnosuke +9 · 12 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- SpoofCeleb: Speech Deepfake Detection and SASV In The Wild
2024/09/18 by Jee-weon Jung, Jung, Jee-weon, Yihan Wu +25 · 24 citations
Computer Science · #Speech Recognition and Synthesis
- PJS: phoneme-balanced Japanese singing voice corpus
2020/06/04 by Junya Koguchi, Shinnosuke Takamichi, Koguchi, Junya +1 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- JVS-MuSiC: Japanese multispeaker singing-voice corpus
2020/01/20 by Hiroki Tamaru, Tamaru, Hiroki, Shinnosuke Takamichi +5 · 7 citations
Computer Science · Medicine · #Speech Recognition and Synthesis #Music and Audio Processing #Voice and Speech Disorders
- Phase reconstruction from amplitude spectrograms based on von-Mises-distribution deep neural network
2018/07/10 by Shinnosuke Takamichi, Yuki Saito, Takamichi, Shinnosuke +7 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- JTubeSpeech: corpus of Japanese speech collected from YouTube for speech recognition and speaker verification
2021/12/17 by Takamichi, Shinnosuke, Kürzinger, Ludwig, Saeki, Takaaki +2 · 6 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- ESPnet2-TTS: Extending the Edge of TTS Research
2021/10/15 by Tomoki Hayashi, Ryuichi Yamamoto, Hayashi, Tomoki +17 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
2024/04/04 by Detai Xin, Xu Tan, Xin, Detai +19 · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Voice Conversion Using Sequence-to-Sequence Learning of Context Posterior Probabilities
2017/04/10 by Hiroyuki Miyoshi, Miyoshi, Hiroyuki, Yuki Saito +5 · 4 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Topic Modeling
- Coco-Nut: Corpus of Japanese Utterance and Voice Characteristics Description for Prompt-based Control
2023/09/24 by Aya Watanabe, Watanabe, Aya, Shinnosuke Takamichi +9 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Text-To-Speech Synthesis In The Wild
2024/09/13 by Jee-weon Jung, Wangyou Zhang, Jung, Jee-weon +25 · 6 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering
- JNV Corpus: A Corpus of Japanese Nonverbal Vocalizations with Diverse Phrases and Emotions
2023/05/21 by Detai Xin, Shinnosuke Takamichi, Xin, Detai +3 · 3 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Learning to Speak from Text: Zero-Shot Multilingual Text-to-Speech with Unsupervised Text Pretraining
2023/01/30 by Takaaki Saeki, Soumi Maiti, Saeki, Takaaki +9 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis
2022/01/26 by Shinnosuke Takamichi, Wataru Nakata, Takamichi, Shinnosuke +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- SelfRemaster: Self-Supervised Speech Restoration with Analysis-by-Synthesis Approach Using Channel Modeling
2022/03/24 by Takaaki Saeki, Saeki, Takaaki, Shinnosuke Takamichi +7 · 2 citations
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis
- Text-to-speech synthesis from dark data with evaluation-in-the-loop data selection
2022/10/26 by Kentaro Seki, Shinnosuke Takamichi, Seki, Kentaro +5 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Improving Speech Prosody of Audiobook Text-to-Speech Synthesis with Acoustic and Textual Contexts
2022/11/04 by Xin, Detai, Adavanne, Sharath, Ang, Federico +3 · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Statistical Parametric Speech Synthesis Incorporating Generative Adversarial Networks
2017/09/23 by Yuki Saito, Saito, Yuki, Shinnosuke Takamichi +3 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- RWCP-SSD-Onomatopoeia: Onomatopoeic Word Dataset for Environmental Sound Synthesis
2020/07/09 by Yuki Okamoto, Okamoto, Yuki, Keisuke Imoto +9 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · #Animal Vocal Communication and Behavior #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- JSSS: free Japanese speech corpus for summarization and simplification
2020/10/05 by Shinnosuke Takamichi, Mamoru Komachi, Takamichi, Shinnosuke +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Text Readability and Simplification #Topic Modeling #electronic engineering #information engineering
- STUDIES: Corpus of Japanese Empathetic Dialogue Speech Towards Friendly Voice Agent
2022/03/28 by Yuki Saito, Saito, Yuki, Yuto Nishimura +7 · 1 citation
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Natural Language Processing Techniques #Social Robot Interaction and HRI #Sound (cs.SD) #Speech and dialogue systems
- Do learned speech symbols follow Zipf's law?
2023/09/18 by Shinnosuke Takamichi, Hiroki Maeda, Takamichi, Shinnosuke +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- CALLS: Japanese Empathetic Dialogue Speech Corpus of Complaint Handling and Attentive Listening in Customer Center
2023/05/23 by Saito, Yuki, Iimori, Eiji, Takamichi, Shinnosuke +2 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Improving robustness of spontaneous speech synthesis with linguistic speech regularization and pseudo-filled-pause insertion
2022/10/18 by Yuta Matsunaga, Takaaki Saeki, Matsunaga, Yuta +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue History
2022/06/16 by Nishimura, Yuto, Saito, Yuki, Takamichi, Shinnosuke +2 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Mid-attribute speaker generation using optimal-transport-based interpolation of Gaussian mixture models
2022/10/18 by Aya Watanabe, Shinnosuke Takamichi, Watanabe, Aya +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus
2023/05/21 by Xin, Detai, Takamichi, Shinnosuke, Morimatsu, Ai +1 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Diversity-based core-set selection for text-to-speech with linguistic and acoustic features
2023/09/15 by Kentaro Seki, Seki, Kentaro, Shinnosuke Takamichi +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
2024/07/22 by Nakata, Wataru, Seki, Kentaro, Yanaka, Hitomi +3 · 3 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- V2S attack: building DNN-based voice conversion from automatic speaker verification
2019/08/05 by Taiki Nakamura, Nakamura, Taiki, Yuki Saito +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
2024/07/05 by Hitoshi Suda, Aya Watanabe, Suda, Hitoshi +3 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
- SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
2024/06/11 by Saito, Yuki, Igarashi, Takuto, Seki, Kentaro +4 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- How Should We Evaluate Synthesized Environmental Sounds
2022/08/16 by Yuki Okamoto, Keisuke Imoto, Okamoto, Yuki +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Sound (cs.SD) #electronic engineering #information engineering
- Building speech corpus with diverse voice characteristics for its prompt-based representation
2024/03/20 by Aya Watanabe, Shinnosuke Takamichi, Watanabe, Aya +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Environmental sound synthesis from vocal imitations and sound event labels
2023/04/29 by Okamoto, Yuki, Imoto, Keisuke, Takamichi, Shinnosuke +3 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
2024/06/11 by T. IGARASHI, Igarashi, Takuto, Yuki Saito +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- TTSOps: A Closed-Loop Corpus Optimization Framework for Training Multi-Speaker TTS Models from Dark Data
2025/06/18 by Kentaro Seki, Seki, Kentaro, Shinnosuke Takamichi +5 · 4 citations
Computer Science · #FOS: Computer and information sciences #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling
- DNN-based Speaker Embedding Using Subjective Inter-speaker Similarity for Multi-speaker Modeling in Speech Synthesis
2019/07/19 by Yuki Saito, Saito, Yuki, Shinnosuke Takamichi +3 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing