Eng Siong Chng
- SpEx+: A Complete Time Domain Speaker Extraction Network
2020/05/10 by Meng Ge, Ge, Meng, Chenglin Xu +9 · 21 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
2025/05/19 by Ziyang Ma, Yinghao Ma, Ma, Ziyang +62 · 53 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- L-SpEx: Localized Target Speaker Extraction
2022/02/21 by Meng Ge, Ge, Meng, Chenglin Xu +9 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
2024/06/02 by Chen Chen, Chen, Chen, Yuchen Hu +9 · 15 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Robotics and Automated Systems #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Speech Emotion Recognition with Co-Attention based Multi-level Acoustic Information
2022/03/29 by Heqing Zou, Zou, Heqing, Yuke Si +7 · 7 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
2025/02/05 by Jixun Yao, Yao, Jixun, Hexin Liu +9 · 20 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- Multi-stage Speaker Extraction with Utterance and Frame-Level Reference Signals
2020/11/19 by Meng Ge, Ge, Meng, Chenglin Xu +9 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
2025/01/27 by Chen Chen, Chen, Chen, Yuchen Hu +13 · 16 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
2024/07/02 by Yuchen Hu, Chen Chen, Hu, Yuchen +7 · 9 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Noise-robust Speech Recognition with 10 Minutes Unparalleled In-domain Data
2022/03/29 by Chen Chen, Nana Hou, Chen, Chen +7 · 5 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
2024/05/16 by Yuchen Hu, Hu, Yuchen, Chen Chen +9 · 8 citations
Computer Science · #Speech Recognition and Synthesis
- LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
2024/09/23 by Hieu-Thi Luong, Haoyang Li, Luong, Hieu-Thi +7 · 10 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
2024/02/10 by Yuchen Hu, Chen Chen, Hu, Yuchen +11 · 6 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
- Adapting OpenAI's Whisper for Speech Recognition on Code-Switch Mandarin-English SEAME and ASRU2019 Datasets
2023/11/29 by Yuhang Yang, Yang, Yuhang, Yizhou Peng +7 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition
2023/05/16 by Yu‐Chen Hu, Ruizhe Li, Hu, Yuchen +9 · 4 citations
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Rainbow Keywords: Efficient Incremental Learning for Online Spoken Keyword Spotting
2022/03/30 by Xiao Yang, Xiao, Yang, Nana Hou +3 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model
2024/02/16 by Xiangyu Zhang, Daijiao Liu, Zhang, Xiangyu +13 · 6 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech and Audio Processing #electronic engineering #information engineering
- Towards Audio Codec-based Speech Separation
2024/06/18 by Jia Qi Yip, Shengkui Zhao, Yip, Jia Qi +7 · 6 citations
Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- NPU-NTU System for Voice Privacy 2024 Challenge
2024/09/06 by Jixun Yao, Nikita Kuzmin, Yao, Jixun +15 · 6 citations
Engineering · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Robotics and Automated Systems #electronic engineering #information engineering
- I2CR: Improving Noise Robustness on Keyword Spotting Using Inter-Intra Contrastive Regularization
2022/09/14 by Dianwen Ng, Ng, Dianwen, Jia Qi Yip +15 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models
2024/05/23 by Yuchen Hu, Chen Chen, Hu, Yuchen +11 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Minimum word error training for non-autoregressive Transformer-based code-switching ASR
2021/10/07 by Yizhou Peng, Jicheng Zhang, Peng, Yizhou +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Wireless Signal Modulation Classification #electronic engineering #information engineering
- Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
2024/09/15 by Chao-Han Huck Yang, Taejin Park, Yang, Chao-Han Huck +39 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Continual Learning For On-Device Environmental Sound Classification
2022/07/15 by Xiao Yang, Xubo Liu, Xiao, Yang +11 · 2 citations
Computer Science · Earth and Planetary Sciences · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
- Speech Separation using Neural Audio Codecs with Embedding Loss
2024/11/27 by Jia Qi Yip, Yip, Jia Qi, Chin Yuen Kwok +5 · 4 citations
Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Generative error correction for code-switching speech recognition using large language models
2023/10/17 by Chen Chen, Chen, Chen, Yuchen Hu +9 · 2 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems
- FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
2025/07/25 by Y Peng, Yi Chao, Peng, Yizhou +11 · 5 citations
Computer Science · Decision Sciences · #Speech and dialogue systems #Personal Information Management and User Behavior #Topic Modeling
- Adapting BERT for Word Sense Disambiguation with Gloss Selection\n Objective and Example Sentences
2020/09/24 by Boon Peng Yap, Andrew Y. Koh, Yap, Boon Peng +3 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- Distilling a speech and music encoder with task arithmetic
2025/05/19 by Fabian Ritter-Gutierrez, Ritter-Gutierrez, Fabian, Yi‐Cheng Lin +10 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Interactive Audio-text Representation for Automated Audio Captioning with Contrastive Learning
2022/03/29 by Chen Chen, Nana Hou, Chen, Chen +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Dataset-Distillation Generative Model for Speech Emotion Recognition
2024/06/05 by Fabian Ritter-Gutierrez, Kuan-Po Huang, Ritter-Gutierrez, Fabian +11 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Advanced Data Compression Techniques
- Self-critical Sequence Training for Automatic Speech Recognition
2022/04/13 by Chen Chen, Chen, Chen, Yu‐Chen Hu +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Intermediate-layer output Regularization for Attention-based Speech Recognition with Shared Decoder
2022/07/09 by Jicheng Zhang, Yizhou Peng, Zhang, Jicheng +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- A Vocoder-free WaveNet Voice Conversion with Non-Parallel Data
2019/02/11 by Xiaohai Tian, Tian, Xiaohai, Eng Siong Chng +3 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge
2024/01/07 by He Wang, Pengcheng Guo, Wang, He +29 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech Separation
2023/02/22 by Yu‐Chen Hu, Hu, Yuchen, Chen Chen +7 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Leveraging Audio-Tagging Assisted Sound Event Detection using Weakified Strong Labels and Frequency Dynamic Convolutions
2023/04/25 by Tanmay Khandelwal, Khandelwal, Tanmay, Rohan Kumar Das +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Speech and Audio Processing #electronic engineering #information engineering
- UniS-MMC: Multimodal Classification via Unimodality-supervised Multimodal Contrastive Learning
2023/05/16 by Heqing Zou, Zou, Heqing, Meng Shen +9 · 1 citation
Computer Science · Engineering · #Advanced Chemical Sensor Technologies #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Text and Document Classification Technologies
- Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding
2025/05/12 by Dianwen Ng, Kun Zhou, Ng, Dianwen +9 · 3 citations
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- SPGM: Prioritizing Local Features for enhanced speech separation performance
2023/09/22 by Jia Qi Yip, Yip, Jia Qi, Shengkui Zhao +19 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Noise-aware Speech Enhancement using Diffusion Probabilistic Model
2023/07/16 by Yu‐Chen Hu, Chen Chen, Hu, Yuchen +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Are Soft Prompts Good Zero-shot Learners for Speech Recognition?
2023/09/18 by Dianwen Ng, Chong Zhang, Ng, Dianwen +17 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Independent language modeling architecture for end-to-end ASR
2019/11/25 by Van Tung Pham, Pham, Van Tung, Haihua Xu +13 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
- Room Impulse Responses help attackers to evade Deep Fake Detection
2024/09/23 by Hieu-Thi Luong, Luong, Hieu-Thi, Duc-Tuan Truong +5 · 2 citations
Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
2025/11/02 by Yayue Deng, Guoqiang Hu, Deng, Yayue +15 · 3 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- The ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge (ICSRC): Dataset, Tracks, Baseline and Results
2022/11/03 by Ao Zhang, Fan Yu, Zhang, Ao +17 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and dialogue systems
- Noise robust distillation of self-supervised speech models via correlation metrics
2023/12/19 by Fabian Ritter-Gutierrez, Kuan-Po Huang, Ritter-Gutierrez, Fabian +11 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Internal Language Model Estimation based Language Model Fusion for Cross-Domain Code-Switching Speech Recognition
2022/07/09 by Yizhou Peng, Peng, Yizhou, Yufei Liu +11 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models
2025/10/10 by Donghang Wu, Wu, Donghang, Haoyang Zhang +20 · 4 citations
Computer Science · Psychology · #Action Observation and Synchronization #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
- Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
2025/09/29 by Hexin Liu, Haoyang Zhang, Liu, Hexin +11 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing