vix.ing · top · new · best · stats · spec

Kashiwagi, Yosuke

  1. Phoneme-aware Encoding for Prefix-tree-based Contextual ASR
    2023/12/15 by Hayato Futami, Emiru Tsunoo, Futami, Hayato +9 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  2. Transformer ASR with Contextual Block Processing
    2019/10/16 by Emiru Tsunoo, Tsunoo, Emiru, Yosuke Kashiwagi +5 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
  3. UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
    2023/10/04 by Siddhant Arora, Hayato Futami, Arora, Siddhant +13 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  4. Gaussian Kernelized Self-Attention for Long Sequence Data and Its Application to CTC-based Speech Recognition
    2021/02/18 by Kashiwagi, Yosuke, Tsunoo, Emiru, Watanabe, Shinji · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  5. Improving Character Error Rate Is Not Equal to Having Clean Speech: Speech Enhancement for ASR Systems with Black-box Acoustic Models
    2021/10/12 by Sawata, Ryosuke, Kashiwagi, Yosuke, Takahashi, Shusuke · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  6. ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
    2025/03/11 by Siddhant Arora, Yifan Peng, Arora, Siddhant +21 · 1 voice · 4 citations
    Computer Science · #Speech and dialogue systems #Multi-Agent Systems and Negotiation #Natural Language Processing Techniques
  7. Streaming Joint Speech Recognition and Disfluency Detection
    2022/11/16 by Hayato Futami, Emiru Tsunoo, Futami, Hayato +11 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
    2025/05/31 by Siddhant Arora, Arora, Siddhant, Jinchuan Tian +13 · 5 citations
    Computer Science · #Speech and dialogue systems #Intelligent Tutoring Systems and Adaptive Learning #Topic Modeling
  9. Tensor decomposition for minimization of E2E SLU model toward on-device processing
    2023/06/02 by Kashiwagi, Yosuke, Arora, Siddhant, Futami, Hayato +6 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  10. Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
    2023/09/16 by Emiru Tsunoo, Tsunoo, Emiru, Hayato Futami +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  11. Decoder-only Architecture for Streaming End-to-end Speech Recognition
    2024/06/23 by Emiru Tsunoo, Tsunoo, Emiru, Hayato Futami +7 · 2 citations
    Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data
    2025/06/02 by Yosuke Kashiwagi, Kashiwagi, Yosuke, Hayato Futami +5 · 2 citations
    Computer Science · #Speech Recognition and Synthesis