vix.ing · top · new · best · stats · spec

Chao-Han Huck Yang

  1. Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
    2023/12/23 by Guan-Ting Lin, Lin, Guan-Ting, Prashanth Gurunath Shivakumar +15 · 14 citations
    Computer Science · #Topic Modeling #Sentiment Analysis and Opinion Mining #Natural Language Processing Techniques
  2. Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
    2024/01/19 by Yuchen Hu, Hu, Yuchen, Chen Chen +10 · 9 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
  3. DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
    2024/09/30 by Ke-Han Lu, Zhehuai Chen, Lu, Ke-Han +13 · 11 citations
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
  4. When BERT Meets Quantum Temporal Convolution Learning for Text Classification in Heterogeneous Computing
    2022/02/17 by Chao-Han Huck Yang, Jun Qi, Yang, Chao-Han Huck +7 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Distributed #FOS: Computer and information sciences #FOS: Electrical engineering #Neural Networks and Reservoir Computing #Neural and Evolutionary Computing (cs.NE) #Parallel #Quantum Computing Algorithms and Architecture #Stochastic Gradient Optimization Techniques #and Cluster Computing (cs.DC) #electronic engineering #information engineering
  5. Towards Neural Scaling Laws for Time Series Foundation Models
    2024/10/16 by Chao-Han Huck Yang, Yao, Qingren, Renhe Jiang +8 · 10 citations
    Computer Science · #Neural Networks and Applications
  6. Chain-of-Thought Prompting for Speech Translation
    2024/09/17 by Ke Hu, Hu, Ke, Zhehuai Chen +13 · 6 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
  7. Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
    2025/01/27 by Chen Chen, Chen, Chen, Yuchen Hu +13 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
    2025/02/14 by William Chen, Chen, William, Jinchuan Tian +9 · 9 citations
    Computer Science · #Speech Recognition and Synthesis
  9. Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
    2024/09/23 by Yuanchao Li, Yuan Gong, Li, Yuanchao +7 · 5 citations
    Psychology · Computer Science · #Emotion and Mood Recognition #Sentiment Analysis and Opinion Mining #Anomaly Detection Techniques and Applications
  10. Low-Resource Music Genre Classification with Cross-Modal Neural Model Reprogramming
    2022/11/02 by Yun-Ning Hung, Hung, Yun-Ning, Chao-Han Huck Yang +5 · 2 citations
    Computer Science · Arts and Humanities · #Music and Audio Processing #Music Technology and Sound Studies #Diverse Musicological Studies
  11. A Lottery Ticket Hypothesis Framework for Low-Complexity Device-Robust Neural Acoustic Scene Classification
    2021/07/03 by Chao-Han Huck Yang, Yen, Hao, Yang, Chao-Han Huck +20 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. ESPnet-SpeechLM: An Open Speech Language Model Toolkit
    2025/02/21 by Jinchuan Tian, Tian, Jinchuan, Jiatong Shi +29 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  13. UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
    2025/03/02 by Alexander H. Liu, Sang-gil Lee, Liu, Alexander H. +13 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  14. Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models
    2024/05/23 by Yuchen Hu, Chen Chen, Hu, Yuchen +11 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  15. Bayesian Example Selection Improves In-Context Learning for Speech, Text, and Visual Modalities
    2024/04/23 by Siyin Wang, Chao-Han Huck Yang, Wang, Siyin +5 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  16. NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
    2024/11/08 by Yen‐Ting Lin, Zhehuai Chen, Lin, Yen-Ting +23 · 2 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques
  17. Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
    2024/09/15 by Chao-Han Huck Yang, Taejin Park, Yang, Chao-Han Huck +39 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  18. Voice Memory for Agentic Speech Recognition
    2026/07/29 by Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko +3
    Computer Science · Engineering · #cs.AI #cs.CL #cs.SD #eess.AS