vix.ing · top · new · best · stats · spec

Kawahara, Tatsuya

  1. Multilingual Turn-taking Prediction Using Voice Activity Projection
    2024/03/11 by Koji Inoue, Inoue, Koji, Bing’er Jiang +7 · 9 citations
    Computer Science · #Speech and dialogue systems
  2. Orthros: Non-autoregressive End-to-end Speech Translation with Dual-decoder
    2020/10/25 by Hirofumi Inaguma, Inaguma, Hirofumi, Yosuke Higuchi +7 · 4 citations
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  3. Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
    2024/01/10 by Koji Inoue, Bing’er Jiang, Inoue, Koji +7 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  4. Zero- and Few-shot Sound Event Localization and Detection
    2023/09/17 by Kazuki Shimada, Shimada, Kazuki, Kengo Uchida +11 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. Distilling the Knowledge of BERT for CTC-based ASR
    2022/09/05 by Futami, Hayato, Inaguma, Hirofumi, Mimura, Masato +2 · 3 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  6. Distilling the Knowledge of BERT for Sequence-to-Sequence ASR
    2020/08/09 by Hayato Futami, Hirofumi Inaguma, Futami, Hayato +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  7. Alignment Knowledge Distillation for Online Streaming Attention-based Speech Recognition
    2021/02/28 by Hirofumi Inaguma, Tatsuya Kawahara, Inaguma, Hirofumi +1 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  8. Time-domain Speech Enhancement Assisted by Multi-resolution Frequency Encoder and Decoder
    2023/03/26 by Hao Shi, Shi, Hao, Masato Mimura +7 · 3 citations
    Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. Non-autoregressive Error Correction for CTC-based ASR with Phone-conditioned Masked LM
    2022/09/08 by Hayato Futami, Futami, Hayato, Hirofumi Inaguma +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  10. Speech Corpus of Ainu Folklore and End-to-end Speech Recognition for Ainu Language
    2020/02/16 by Kohei Matsuura, Sei Ueno, Matsuura, Kohei +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  11. Multilingual End-to-End Speech Translation
    2019/10/01 by Hirofumi Inaguma, Kevin Duh, Inaguma, Hirofumi +5 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  12. MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score Prediction
    2024/01/24 by Zhou, Wangjin, Yang, Zhengdong, Chu, Chenhui +4 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #electronic engineering #information engineering
  13. Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
    2024/10/21 by Koji Inoue, Inoue, Koji, Divesh Lala +5 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  14. Designing Precise and Robust Dialogue Response Evaluators
    2020/04/10 by Tianyu Zhao, Zhao, Tianyu, Divesh Lala +3 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  15. Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
    2023/05/18 by Shi, Hao, Shimada, Kazuki, Hirano, Masato +6 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  16. Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
    2024/09/01 by Hao Shi, Yuan Gao, Shi, Hao +5 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  17. Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring
    2021/09/09 by Hirofumi Inaguma, Inaguma, Hirofumi, Yosuke Higuchi +7 · 1 citation
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech Recognition and Synthesis
  18. Intelligent Conversational Android ERICA Applied to Attentive Listening and Job Interview
    2021/05/02 by Tatsuya Kawahara, Kawahara, Tatsuya, Koji Inoue +3 · 1 citation
    Arts and Humanities · Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Language, Discourse, Communication Strategies #Robotics (cs.RO) #Social Robot Interaction and HRI #Speech and dialogue systems
  19. Reasoning before Responding: Integrating Commonsense-based Causality Explanation for Empathetic Response Generation
    2023/07/28 by Yahui Fu, Koji Inoue, Fu, Yahui +5 · 2 citations
    Computer Science · Medicine · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Simulation-Based Education in Healthcare #Topic Modeling
  20. End-to-end Speech-to-Punctuated-Text Recognition
    2022/07/07 by Jumon Nozaki, Tatsuya Kawahara, Nozaki, Jumon +5 · 1 citation
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  21. Enhancing Long-term RAG Chatbots with Psychological Models of Memory Importance and Forgetting
    2024/09/19 by Ryuichi Sumida, Koji Inoue, Sumida, Ryuichi +3 · 2 citations
    Computer Science · Medicine · #AI in Service Interactions #Artificial Intelligence in Healthcare and Education
  22. Alzheimer's Dementia Detection through Spontaneous Dialogue with Proactive Robotic Listeners
    2022/11/15 by Li, Yuanchao, Lai, Catherine, Lala, Divesh +2 · 1 citation
    #FOS: Computer and information sciences #Robotics (cs.RO)
  23. An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue
    2025/01/28 by Koji Inoue, Inoue, Koji, Divesh Lala +7 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  24. Human-Like Embodied AI Interviewer: Employing Android ERICA in Real\n International Conference
    2024/12/13 by Zi Haur Pang, Pang, Zi Haur, Yahui Fu +9 · 1 citation
    Engineering · #Robotics and Automated Systems
  25. Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
    2024/08/29 by Yuka Ko, Sheng Li, Ko, Yuka +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  26. Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
    2024/09/12 by Wangjin Zhou, Zhou, Wangjin, Fengrun Zhang +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  27. Exploration of Adapter for Noise Robust Automatic Speech Recognition
    2024/02/28 by Hao Shi, Tatsuya Kawahara, Shi, Hao +1 · 1 citation
    Computer Science · #Speech Recognition and Synthesis