Hung-yi Lee
- SUPERB: Speech processing Universal PERformance Benchmark
2021/05/03 by Shu-Wen Yang, Yang, Shu-wen, Po-Han Chi +37 · 69 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Can Large Language Models Be an Alternative to Human Evaluations?
2023/05/03 by Cheng-Han Chiang, Hung-yi Lee, Chiang, Cheng-Han +1 · 90 citations
Computer Science · #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Natural Language Processing Techniques #Topic Modeling
- Temporal Pattern Attention for Multivariate Time Series Forecasting
2018/09/12 by Shun-Yao Shih, Fan-Keng Sun, Shih, Shun-Yao +3 · 11 citations
Computer Science · Decision Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Music and Audio Processing #Stock Market Forecasting Methods #Time Series Analysis and Forecasting
- Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
2023/09/18 by Chien‐Yu Huang, Huang, Chien-yu, Ke-Han Lu +27 · 20 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden-unit BERT
2021/10/05 by Heng-Jui Chang, Shuwen Yang, Chang, Heng-Jui +3 · 11 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- LAMOL: LAnguage MOdeling for Lifelong Language Learning
2019/09/07 by Fan-Keng Sun, Cheng-Hao Ho, Sun, Fan-Keng +3 · 8 citations
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Topic Modeling
- On The Landscape of Spoken Language Models: A Comprehensive Survey
2025/04/11 by Siddhant Arora, Arora, Siddhant, Kai‐Wei Chang +17 · 37 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations
2024/02/20 by Guan-Ting Lin, Cheng-Han Chiang, Lin, Guan-Ting +3 · 16 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech and dialogue systems #electronic engineering #information engineering
- Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
2023/12/23 by Guan-Ting Lin, Lin, Guan-Ting, Prashanth Gurunath Shivakumar +15 · 14 citations
Computer Science · #Topic Modeling #Sentiment Analysis and Opinion Mining #Natural Language Processing Techniques
- Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models
2024/08/05 by Zhi Rui Tam, Cheng-Kuang Wu, Tam, Zhi Rui +9 · 16 citations
Computer Science · #Topic Modeling
- AGAIN-VC: A One-shot Voice Conversion using Activation Guidance and Adaptive Instance Normalization
2020/10/31 by Yen-Hao Chen, Da-Yi Wu, Chen, Yen-Hao +5 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Towards audio language modeling -- an overview
2024/02/20 by Haibin Wu, Wu, Haibin, Xuanjun Chen +11 · 13 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset
2024/09/13 by Jiawei Du, Lin I, Du, Jiawei +14 · 15 citations
Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
2024/09/30 by Ke-Han Lu, Zhehuai Chen, Lu, Ke-Han +13 · 15 citations
Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
- Over-Reasoning and Redundant Calculation of Large Language Models
2024/01/21 by Cheng-Han Chiang, Chiang, Cheng-Han, Hung-yi Lee +1 · 9 citations
Computer Science · Decision Sciences · #Topic Modeling #Natural Language Processing Techniques #Data Quality and Management
- StreamBench: Towards Benchmarking Continuous Improvement of Language Agents
2024/06/13 by Cheng-Kuang Wu, Zhi Rui Tam, Wu, Cheng-Kuang +7 · 9 citations
Computer Science · #Multi-Agent Systems and Negotiation #Semantic Web and Ontologies #Natural Language Processing Techniques
- On the Utility of Self-supervised Models for Prosody-related Tasks
2022/10/13 by Guan-Ting Lin, Lin, Guan-Ting, Chi-Luen Feng +13 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Tree Transformer: Integrating Tree Structures into Self-Attention
2019/09/14 by Yau-Shian Wang, Hung-yi Lee, Wang, Yau-Shian +3 · 4 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
2024/06/27 by Ke-Han Lu, Zhehuai Chen, Lu, Ke-Han +11 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Code-switching Sentence Generation by Generative Adversarial Networks and its Application to Data Augmentation
2018/11/06 by Ching-Ting Chang, Shun-Po Chuang, Chang, Ching-Ting +3 · 4 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback
2024/11/04 by Guan-Ting Lin, Prashanth Gurunath Shivakumar, Lin, Guan-Ting +11 · 11 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
- Listen, Adapt, Better WER: Source-free Single-utterance Test-time Adaptation for Automatic Speech Recognition
2022/03/27 by Guan-Ting Lin, Shang-Wen Li, Lin, Guan-Ting +3 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
2024/06/11 by Haibin Wu, Yuan Tseng, Wu, Haibin +3 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Worse WER, but Better BLEU? Leveraging Word Embedding as Intermediate in Multitask End-to-End Speech Translation
2020/05/21 by Shun-Po Chuang, Tzu-Wei Sung, Chuang, Shun-Po +5 · 3 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech Recognition and Synthesis
- Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
2024/10/21 by Chun-Yi Kuan, Kuan, Chun-Yi, Hung-yi Lee +1 · 8 citations
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
- The defender's perspective on automatic speaker verification: An overview
2023/05/22 by Haibin Wu, Jiawen Kang, Wu, Haibin +7 · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Network Security and Intrusion Detection #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
2024/06/12 by Chun-Yi Kuan, Kuan, Chun-Yi, Weiping Huang +3 · 6 citations
Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Digital Media Forensic Detection
- Why We Should Report the Details in Subjective Evaluation of TTS More Rigorously
2023/06/03 by Cheng-Han Chiang, Wei‐Ping Huang, Chiang, Cheng-Han +3 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Signal Processing (eess.SP) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Merging Facts, Crafting Fallacies: Evaluating the Contradictory Nature of Aggregated Factual Claims in Long-Form Generations
2024/02/08 by Cheng-Han Chiang, Chiang, Cheng-Han, Hung-yi Lee +1 · 4 citations
Economics, Econometrics and Finance · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Law, Economics, and Judicial Systems
- Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
2024/07/09 by Yi‐Cheng Lin, Lin, Yi-Cheng, Tzu-Quan Lin +11 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #electronic engineering #information engineering
- Discrete Audio Tokens: More Than a Survey!
2025/06/12 by Pooneh Mousavi, Mousavi, Pooneh, Gallil Maimon +39 · 13 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
2024/09/17 by Hsi-Che Lin, Yi‐Cheng Lin, Lin, Hsi-Che +5 · 6 citations
Computer Science · #Speech Recognition and Synthesis
- End-to-end Text-to-speech for Low-resource Languages by Cross-Lingual Transfer Learning
2019/04/13 by Tao Tu, Yuan-Jui Chen, Tu, Tao +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Adversarial Attacks on Spoofing Countermeasures of automatic speaker verification
2019/10/19 by Songxiang Liu, Liu, Songxiang, Haibin Wu +5 · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Network Security and Intrusion Detection #Speech Recognition and Synthesis #electronic engineering #information engineering
- Audio-Aware Large Language Models as Judges for Speaking Styles
2025/06/06 by Cheng-Han Chiang, Xiaofei Wang, Chiang, Cheng-Han +18 · 10 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #electronic engineering #information engineering
- EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
2024/02/20 by Haibin Wu, Wu, Haibin, Huang-Cheng Chou +13 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models
2024/08/07 by Shachi H Kumar, Saurav Sahay, Kumar, Shachi H +15 · 5 citations
Social Sciences · #Artificial Intelligence in Law #Legal Language and Interpretation
- Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models
2025/05/25 by Ke-Han Lu, Chun-Yi Kuan, Lu, Ke-Han +3 · 12 citations
Computer Science · #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- Investigating on Incorporating Pretrained and Learnable Speaker\n Representations for Multi-Speaker Multi-Style Text-to-Speech
2021/03/06 by Chung-Ming Chien, Chien, Chung-Ming, Jheng-Hao Lin +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Towards General-Purpose Text-Instruction-Guided Voice Conversion
2023/09/25 by Chun-Yi Kuan, Chen Li, Kuan, Chun-Yi +13 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- Adversarial Sample Detection for Speaker Verification by Neural Vocoders
2021/07/01 by Haibin Wu, Wu, Haibin, Po‐Chun Hsu +15 · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Meta Learning for Natural Language Processing: A Survey
2022/05/03 by Hung-yi Lee, Lee, Hung-yi, Shang-Wen Li +3 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification #Topic Modeling
- Spoofing-Aware Speaker Verification by Multi-Level Fusion
2022/03/29 by Haibin Wu, Wu, Haibin, Lingwei Meng +13 · 2 citations
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
- Building a Taiwanese Mandarin Spoken Language Model: A First Attempt
2024/11/11 by Chih-Kai Yang, Yang, Chih-Kai, Yu-Kuan Fu +39 · 6 citations
Computer Science · #Natural Language Processing Techniques
- MelHuBERT: A simplified HuBERT on Mel spectrograms
2022/11/17 by Tzu-Quan Lin, Hung-yi Lee, Lin, Tzu-Quan +3 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- Learning to Encode Text as Human-Readable Summaries using Generative Adversarial Networks
2018/10/05 by Yau-Shian Wang, Wang, Yau-Shian, Hung-yi Lee +1 · 2 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
- Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition
2024/06/05 by Hsuan Su, Su, Hsuan, Hua Farn +6 · 3 citations
Engineering · #Fault Detection and Control Systems
- SpeechPrompt v2: Prompt Tuning for Speech Classification Tasks
2023/03/01 by Kai-Wei Chang, Yu–Kai Wang, Chang, Kai-Wei +11 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts
2023/06/03 by Haibin Wu, Kai-Wei Chang, Wu, Haibin +5 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Singing Voice Graph Modeling for SingFake Detection
2024/06/05 by Xuanjun Chen, Haibin Wu, Chen, Xuanjun +5 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- A Closer Look into Automatic Evaluation Using Large Language Models
2023/10/09 by Cheng-Han Chiang, Hung-yi Lee, Chiang, Cheng-Han +1 · 2 citations
Computer Science · Social Sciences · #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks
2023/10/19 by Ming-Hao Hsu, Hsu, Ming-Hao, Kai-Wei Chang +5 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Generative Adversarial Networks for Unpaired Voice Transformation on Impaired Speech
2018/10/30 by Li-Wei Chen, Hung-yi Lee, Chen, Li-Wei +3 · 1 citation
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
- CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech
2025/01/14 by Xuanjun Chen, Jiawei Du, Chen, Xuanjun +19 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Cross-Lingual Transfer Learning for Question Answering
2019/07/13 by Chia‐Hsuan Lee, Lee, Chia-Hsuan, Hung-yi Lee +1 · 1 citation
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
- Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks
2024/01/05 by Kevin Everson, Everson, Kevin, Yile Gu +23 · 2 citations
Computer Science · #Speech and dialogue systems #Topic Modeling #Natural Language Processing Techniques
- Pre-Training a Language Model Without Human Language
2020/12/22 by Cheng-Han Chiang, Hung-yi Lee, Chiang, Cheng-Han +1 · 1 citation
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
- A Preliminary Exploration with GPT-4o Voice Mode
2025/02/14 by Yuxiang Lin, Lin, Yu-Xiang, Chih-Kai Yang +11 · 5 citations
Computer Science · #Computational Physics and Python Applications
- FragmentVC: Any-to-Any Voice Conversion by End-to-End Extracting and\n Fusing Fine-Grained Voice Fragments With Attention
2020/10/27 by Yist Y. Lin, Lin, Yist Y., Chung-Ming Chien +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Analyzing the Robustness of Unsupervised Speech Recognition
2021/10/07 by Guan-Ting Lin, Lin, Guan-Ting, Chan-Jan Hsu +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Don't speak too fast: The impact of data bias on self-supervised speech models
2021/10/15 by Yen Meng, Yi-Hui Chou, Meng, Yen +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- S3PRL-VC: Open-source Voice Conversion Framework with Self-supervised Speech Representations
2021/10/12 by Wen-Chin Huang, Huang, Wen-Chin, Shuwen Yang +9 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
- Dataset-Distillation Generative Model for Speech Emotion Recognition
2024/06/05 by Fabian Ritter-Gutierrez, Kuan-Po Huang, Ritter-Gutierrez, Fabian +11 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Advanced Data Compression Techniques
- On the Evaluation of Speech Foundation Models for Spoken Language Understanding
2024/06/14 by Siddhant Arora, Ankita Pasad, Arora, Siddhant +21 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Bridging Speech and Textual Pre-trained Models with Unsupervised ASR
2022/11/06 by Jiatong Shi, Chan-Jan Hsu, Shi, Jiatong +13 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- DUAL: Discrete Spoken Unit Adaptive Learning for Textless Spoken Question Answering
2022/03/09 by Guan-Ting Lin, Lin, Guan-Ting, Yung-Sung Chuang +17 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling
- Creativity in LLM-based Multi-Agent Systems: A Survey
2025/05/27 by Yi-Cheng Lin, Lin, Yi-Cheng, Kang-Chieh Chen +12 · 3 citations
Computer Science · #Multi-Agent Systems and Negotiation #Semantic Web and Ontologies
- MiniSUPERB: Lightweight Benchmark for Self-supervised Speech Models
2023/05/30 by Yu-Hsiang Wang, Wang, Yu-Hsiang, Huang-Yu Chen +7 · 1 citation
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Improving Textless Spoken Language Understanding with Discrete Units as Intermediate Target
2023/05/29 by Guan-Wei Wu, Wu, Guan-Wei, Guan-Ting Lin +5 · 1 citation
Computer Science · #Topic Modeling #Speech and dialogue systems #Natural Language Processing Techniques
- Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling
2024/12/21 by Shao-Syuan Huang, Huang, Shao-Syuan, Kuan-Po Huang +5 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Speech and dialogue systems
- SUPERB @ SLT 2022: Challenge on Generalization and Efficiency of Self-Supervised Speech Representation Learning
2022/10/16 by Tzu-hsun Feng, Feng, Tzu-hsun, Annie Dong +25 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Mitigating Forgetting in LLM Fine-Tuning via Low-Perplexity Token Learning
2025/01/24 by Chao-Chung Wu, Wu, Chao-Chung, Zhi Rui Tam +8 · 2 citations
Computer Science · #Digital Rights Management and Security
- Can LLMs Understand the Implication of Emphasized Sentences in Dialogue?
2024/06/16 by Guan-Ting Lin, Lin, Guan-Ting, Hung-yi Lee +1 · 2 citations
Health Professions · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Interpreting and Communication in Healthcare
- PEFT for Speech: Unveiling Optimal Placement, Merging Strategies, and Ensemble Techniques
2024/01/04 by Tzu-Han Lin, Lin, Tzu-Han, How-Shing Wang +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
2025/02/18 by Tzu-Quan Lin, Weiping Huang, Lin, Tzu-Quan +5 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
- SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
2024/01/24 by Chyi-Jiunn Lin, Lin, Chyi-Jiunn, Guan-Ting Lin +13 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Transferring Textual Preferences to Vision-Language Understanding through Model Merging
2025/02/19 by Chen-An Li, Tzu-Han Lin, Li, Chen-An +5 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Semantic Web and Ontologies #Speech and dialogue systems
- DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging
2024/07/01 by Tzu-Han Lin, Chen-An Li, Lin, Tzu-Han +5 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
2024/06/09 by Chih-Kai Yang, Yang, Chih-Kai, Kuan-Po Huang +3 · 1 citation
Decision Sciences · #Ethics in Business and Education
- Meta-PerSER: Few-Shot Listener Personalized Speech Emotion Recognition via Meta-learning
2025/05/22 by Shen, Liang-Yeh, Yi‐Cheng Lin, Fang, Shi-Xin +5 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sentiment Analysis and Opinion Mining #Speech Recognition and Synthesis #electronic engineering #information engineering
- Defense for Black-box Attacks on Anti-spoofing Models by Self-Supervised Learning
2020/06/05 by Haibin Wu, Wu, Haibin, Andy T. Liu +3 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #Machine Learning (cs.LG) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples
2025/05/20 by Chun-Yi Kuan, Kuan, Chun-Yi, Hung-yi Lee +1 · 2 citations
Computer Science · Psychology · #Music and Audio Processing #Emotion and Mood Recognition #Explainable Artificial Intelligence (XAI)
- Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond
2023/10/09 by Jiatong Shi, William Chen, Shi, Jiatong +23 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Unveiling Narrative Reasoning Limits of Large Language Models with Trope in Movie Synopses
2024/09/22 by Hung-Ting Su, Su, Hung-Ting, Ya-Ching Hsu +11 · 1 citation
Computer Science · Social Sciences · #Topic Modeling #Natural Language Processing Techniques #Computational and Text Analysis Methods
- Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding
2025/06/08 by T. C. Hsu, Hsu, Tzu-wen, Ke-Han Lu +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling
2025/05/31 by Kuan-Po Huang, Huang, Kuan-Po, Shu-Wen Yang +19 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging
2024/12/27 by Hua Farn, Farn, Hua, Hsuan Su +9 · 1 citation
Engineering · #Semiconductor materials and devices
- Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models
2026/07/26 by Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim +7
#eess.AS #cs.AI #cs.CL #cs.LG #cs.SD
- EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
2026/07/20 by Jia-Kai Dong, Yi-Cheng Lin, Hung-yi Lee
#cs.HC #cs.AI