Yang, Chih-Kai
- Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
2024/11/08 by Huang, Chien-yu, Chen, Wei-Chih, Yang, Shu-wen +77 · 34 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
2024/07/09 by Yi‐Cheng Lin, Lin, Yi-Cheng, Tzu-Quan Lin +11 · 9 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #electronic engineering #information engineering
- Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
2023/10/04 by Kuan-Po Huang, Huang, Kuan-Po, Chih-Kai Yang +7 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
2025/05/19 by Yang, Chih-Kai, Ho, Neo, Piao, Yen-Ting +1 · 17 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey
2025/05/21 by Yang, Chih-Kai, Ho, Neo S., Lee, Hung-yi · 19 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
2025/07/03 by Lu, Ke-Han, Chen, Zhehuai, Fu, Szu-Wei +25 · 19 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Building a Taiwanese Mandarin Spoken Language Model: A First Attempt
2024/11/11 by Chih-Kai Yang, Yu-Kuan Fu, Yang, Chih-Kai +39 · 8 citations
Computer Science · #Natural Language Processing Techniques
- Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
2024/07/13 by Kuan, Chun-Yi, Yang, Chih-Kai, Huang, Wei-Ping +2 · 6 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision
2023/12/30 by Yang, Chih-Kai, Huang, Kuan-Po, Lu, Ke-Han +3 · 4 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- A Preliminary Exploration with GPT-4o Voice Mode
2025/02/14 by Yuxiang Lin, Lin, Yu-Xiang, Chih-Kai Yang +11 · 7 citations
Computer Science · #Computational Physics and Python Applications
- Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
2024/06/09 by Chih-Kai Yang, Kuan-Po Huang, Yang, Chih-Kai +3 · 3 citations
Decision Sciences · #Ethics in Business and Education
- Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
2025/05/23 by Hsiao, Chi-Yuan, Lu, Ke-Han, Chang, Kai-Wei +3 · 3 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering