vix.ing · top · new · best · stats · spec

Lu, Bo-Ru

  1. Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
    2024/09/07 by Junkai Wu, Wu, Junkai, Xulin Fan +11 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  2. Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
    2025/07/14 by Yang, Shu-wen, Kim, Byeonggeun, Huang, Kuan-Po +8 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  3. Does Collaborative Human-LM Dialogue Generation Help Information Extraction from Human Dialogues?
    2023/07/13 by Lu, Bo-Ru, Haduong, Nikita, Lee, Chia-Hsuan +7 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  4. IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling
    2025/05/31 by Kuan-Po Huang, Huang, Kuan-Po, Shu-Wen Yang +19 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering