Lu, Bo-Ru
- Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
2024/09/07 by Junkai Wu, Wu, Junkai, Xulin Fan +11 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
2025/07/14 by Yang, Shu-wen, Kim, Byeonggeun, Huang, Kuan-Po +8 · 4 citations
#Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Does Collaborative Human-LM Dialogue Generation Help Information Extraction from Human Dialogues?
2023/07/13 by Lu, Bo-Ru, Haduong, Nikita, Lee, Chia-Hsuan +7 · 1 citation
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling
2025/05/31 by Kuan-Po Huang, Huang, Kuan-Po, Shu-Wen Yang +19 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering