Kawahara, Tatsuya
- Multilingual Turn-taking Prediction Using Voice Activity Projection
2024/03/11 by Koji Inoue, Inoue, Koji, Bing’er Jiang +7 · 9 citations
Computer Science · #Speech and dialogue systems
- Orthros: Non-autoregressive End-to-end Speech Translation with Dual-decoder
2020/10/25 by Hirofumi Inaguma, Inaguma, Hirofumi, Yosuke Higuchi +7 · 4 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
2024/01/10 by Koji Inoue, Bing’er Jiang, Inoue, Koji +7 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- Zero- and Few-shot Sound Event Localization and Detection
2023/09/17 by Kazuki Shimada, Shimada, Kazuki, Kengo Uchida +11 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Distilling the Knowledge of BERT for CTC-based ASR
2022/09/05 by Futami, Hayato, Inaguma, Hirofumi, Mimura, Masato +2 · 3 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Distilling the Knowledge of BERT for Sequence-to-Sequence ASR
2020/08/09 by Hayato Futami, Hirofumi Inaguma, Futami, Hayato +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Alignment Knowledge Distillation for Online Streaming Attention-based Speech Recognition
2021/02/28 by Hirofumi Inaguma, Tatsuya Kawahara, Inaguma, Hirofumi +1 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Time-domain Speech Enhancement Assisted by Multi-resolution Frequency Encoder and Decoder
2023/03/26 by Hao Shi, Shi, Hao, Masato Mimura +7 · 3 citations
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Non-autoregressive Error Correction for CTC-based ASR with Phone-conditioned Masked LM
2022/09/08 by Hayato Futami, Futami, Hayato, Hirofumi Inaguma +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Speech Corpus of Ainu Folklore and End-to-end Speech Recognition for Ainu Language
2020/02/16 by Kohei Matsuura, Sei Ueno, Matsuura, Kohei +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Multilingual End-to-End Speech Translation
2019/10/01 by Hirofumi Inaguma, Kevin Duh, Inaguma, Hirofumi +5 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score Prediction
2024/01/24 by Zhou, Wangjin, Yang, Zhengdong, Chu, Chenhui +4 · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #electronic engineering #information engineering
- Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
2024/10/21 by Koji Inoue, Inoue, Koji, Divesh Lala +5 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Designing Precise and Robust Dialogue Response Evaluators
2020/04/10 by Tianyu Zhao, Zhao, Tianyu, Divesh Lala +3 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
- Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
2023/05/18 by Shi, Hao, Shimada, Kazuki, Hirano, Masato +6 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
2024/09/01 by Hao Shi, Yuan Gao, Shi, Hao +5 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring
2021/09/09 by Hirofumi Inaguma, Inaguma, Hirofumi, Yosuke Higuchi +7 · 1 citation
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech Recognition and Synthesis
- Intelligent Conversational Android ERICA Applied to Attentive Listening and Job Interview
2021/05/02 by Tatsuya Kawahara, Kawahara, Tatsuya, Koji Inoue +3 · 1 citation
Arts and Humanities · Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Language, Discourse, Communication Strategies #Robotics (cs.RO) #Social Robot Interaction and HRI #Speech and dialogue systems
- Reasoning before Responding: Integrating Commonsense-based Causality Explanation for Empathetic Response Generation
2023/07/28 by Yahui Fu, Koji Inoue, Fu, Yahui +5 · 2 citations
Computer Science · Medicine · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Simulation-Based Education in Healthcare #Topic Modeling
- End-to-end Speech-to-Punctuated-Text Recognition
2022/07/07 by Jumon Nozaki, Tatsuya Kawahara, Nozaki, Jumon +5 · 1 citation
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Enhancing Long-term RAG Chatbots with Psychological Models of Memory Importance and Forgetting
2024/09/19 by Ryuichi Sumida, Koji Inoue, Sumida, Ryuichi +3 · 2 citations
Computer Science · Medicine · #AI in Service Interactions #Artificial Intelligence in Healthcare and Education
- Alzheimer's Dementia Detection through Spontaneous Dialogue with Proactive Robotic Listeners
2022/11/15 by Li, Yuanchao, Lai, Catherine, Lala, Divesh +2 · 1 citation
#FOS: Computer and information sciences #Robotics (cs.RO)
- An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue
2025/01/28 by Koji Inoue, Inoue, Koji, Divesh Lala +7 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
- Human-Like Embodied AI Interviewer: Employing Android ERICA in Real\n International Conference
2024/12/13 by Zi Haur Pang, Pang, Zi Haur, Yahui Fu +9 · 1 citation
Engineering · #Robotics and Automated Systems
- Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
2024/08/29 by Yuka Ko, Sheng Li, Ko, Yuka +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
2024/09/12 by Wangjin Zhou, Zhou, Wangjin, Fengrun Zhang +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Exploration of Adapter for Noise Robust Automatic Speech Recognition
2024/02/28 by Hao Shi, Tatsuya Kawahara, Shi, Hao +1 · 1 citation
Computer Science · #Speech Recognition and Synthesis