vix.ing · top · new · best · stats · spec

Eng Siong Chng

  1. SpEx+: A Complete Time Domain Speaker Extraction Network
    2020/05/10 by Meng Ge, Ge, Meng, Chenglin Xu +9 · 21 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  2. MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
    2025/05/19 by Ziyang Ma, Yinghao Ma, Ma, Ziyang +62 · 53 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  3. L-SpEx: Localized Target Speaker Extraction
    2022/02/21 by Meng Ge, Ge, Meng, Chenglin Xu +9 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  4. Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
    2024/06/02 by Chen Chen, Chen, Chen, Yuchen Hu +9 · 15 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Robotics and Automated Systems #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  5. Speech Emotion Recognition with Co-Attention based Multi-level Acoustic Information
    2022/03/29 by Heqing Zou, Zou, Heqing, Yuke Si +7 · 7 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  6. GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
    2025/02/05 by Jixun Yao, Yao, Jixun, Hexin Liu +9 · 20 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  7. Multi-stage Speaker Extraction with Utterance and Frame-Level Reference Signals
    2020/11/19 by Meng Ge, Ge, Meng, Chenglin Xu +9 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
    2025/01/27 by Chen Chen, Chen, Chen, Yuchen Hu +13 · 16 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
    2024/07/02 by Yuchen Hu, Chen Chen, Hu, Yuchen +7 · 9 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  10. Noise-robust Speech Recognition with 10 Minutes Unparalleled In-domain Data
    2022/03/29 by Chen Chen, Nana Hou, Chen, Chen +7 · 5 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  11. Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
    2024/05/16 by Yuchen Hu, Hu, Yuchen, Chen Chen +9 · 8 citations
    Computer Science · #Speech Recognition and Synthesis
  12. LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
    2024/09/23 by Hieu-Thi Luong, Haoyang Li, Luong, Hieu-Thi +7 · 10 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  13. GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
    2024/02/10 by Yuchen Hu, Chen Chen, Hu, Yuchen +11 · 6 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
  14. Adapting OpenAI's Whisper for Speech Recognition on Code-Switch Mandarin-English SEAME and ASRU2019 Datasets
    2023/11/29 by Yuhang Yang, Yang, Yuhang, Yizhou Peng +7 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  15. Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition
    2023/05/16 by Yu‐Chen Hu, Ruizhe Li, Hu, Yuchen +9 · 4 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  16. Rainbow Keywords: Efficient Incremental Learning for Online Spoken Keyword Spotting
    2022/03/30 by Xiao Yang, Xiao, Yang, Nana Hou +3 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  17. Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model
    2024/02/16 by Xiangyu Zhang, Daijiao Liu, Zhang, Xiangyu +13 · 6 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech and Audio Processing #electronic engineering #information engineering
  18. Towards Audio Codec-based Speech Separation
    2024/06/18 by Jia Qi Yip, Shengkui Zhao, Yip, Jia Qi +7 · 6 citations
    Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  19. NPU-NTU System for Voice Privacy 2024 Challenge
    2024/09/06 by Jixun Yao, Nikita Kuzmin, Yao, Jixun +15 · 6 citations
    Engineering · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Robotics and Automated Systems #electronic engineering #information engineering
  20. I2CR: Improving Noise Robustness on Keyword Spotting Using Inter-Intra Contrastive Regularization
    2022/09/14 by Dianwen Ng, Ng, Dianwen, Jia Qi Yip +15 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  21. Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models
    2024/05/23 by Yuchen Hu, Chen Chen, Hu, Yuchen +11 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  22. Minimum word error training for non-autoregressive Transformer-based code-switching ASR
    2021/10/07 by Yizhou Peng, Jicheng Zhang, Peng, Yizhou +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Wireless Signal Modulation Classification #electronic engineering #information engineering
  23. Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
    2024/09/15 by Chao-Han Huck Yang, Taejin Park, Yang, Chao-Han Huck +39 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  24. Continual Learning For On-Device Environmental Sound Classification
    2022/07/15 by Xiao Yang, Xubo Liu, Xiao, Yang +11 · 2 citations
    Computer Science · Earth and Planetary Sciences · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
  25. Speech Separation using Neural Audio Codecs with Embedding Loss
    2024/11/27 by Jia Qi Yip, Yip, Jia Qi, Chin Yuen Kwok +5 · 4 citations
    Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  26. Generative error correction for code-switching speech recognition using large language models
    2023/10/17 by Chen Chen, Chen, Chen, Yuchen Hu +9 · 2 citations
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems
  27. FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
    2025/07/25 by Y Peng, Yi Chao, Peng, Yizhou +11 · 5 citations
    Computer Science · Decision Sciences · #Speech and dialogue systems #Personal Information Management and User Behavior #Topic Modeling
  28. Adapting BERT for Word Sense Disambiguation with Gloss Selection\n Objective and Example Sentences
    2020/09/24 by Boon Peng Yap, Andrew Y. Koh, Yap, Boon Peng +3 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  29. Distilling a speech and music encoder with task arithmetic
    2025/05/19 by Fabian Ritter-Gutierrez, Ritter-Gutierrez, Fabian, Yi‐Cheng Lin +10 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  30. Interactive Audio-text Representation for Automated Audio Captioning with Contrastive Learning
    2022/03/29 by Chen Chen, Nana Hou, Chen, Chen +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  31. Dataset-Distillation Generative Model for Speech Emotion Recognition
    2024/06/05 by Fabian Ritter-Gutierrez, Kuan-Po Huang, Ritter-Gutierrez, Fabian +11 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Advanced Data Compression Techniques
  32. Self-critical Sequence Training for Automatic Speech Recognition
    2022/04/13 by Chen Chen, Chen, Chen, Yu‐Chen Hu +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  33. Intermediate-layer output Regularization for Attention-based Speech Recognition with Shared Decoder
    2022/07/09 by Jicheng Zhang, Yizhou Peng, Zhang, Jicheng +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  34. A Vocoder-free WaveNet Voice Conversion with Non-Parallel Data
    2019/02/11 by Xiaohai Tian, Tian, Xiaohai, Eng Siong Chng +3 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  35. ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge
    2024/01/07 by He Wang, Pengcheng Guo, Wang, He +29 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  36. Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech Separation
    2023/02/22 by Yu‐Chen Hu, Hu, Yuchen, Chen Chen +7 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  37. Leveraging Audio-Tagging Assisted Sound Event Detection using Weakified Strong Labels and Frequency Dynamic Convolutions
    2023/04/25 by Tanmay Khandelwal, Khandelwal, Tanmay, Rohan Kumar Das +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Speech and Audio Processing #electronic engineering #information engineering
  38. UniS-MMC: Multimodal Classification via Unimodality-supervised Multimodal Contrastive Learning
    2023/05/16 by Heqing Zou, Zou, Heqing, Meng Shen +9 · 1 citation
    Computer Science · Engineering · #Advanced Chemical Sensor Technologies #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Text and Document Classification Technologies
  39. Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding
    2025/05/12 by Dianwen Ng, Kun Zhou, Ng, Dianwen +9 · 3 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  40. SPGM: Prioritizing Local Features for enhanced speech separation performance
    2023/09/22 by Jia Qi Yip, Yip, Jia Qi, Shengkui Zhao +19 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  41. Noise-aware Speech Enhancement using Diffusion Probabilistic Model
    2023/07/16 by Yu‐Chen Hu, Chen Chen, Hu, Yuchen +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  42. Are Soft Prompts Good Zero-shot Learners for Speech Recognition?
    2023/09/18 by Dianwen Ng, Chong Zhang, Ng, Dianwen +17 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  43. Independent language modeling architecture for end-to-end ASR
    2019/11/25 by Van Tung Pham, Pham, Van Tung, Haihua Xu +13 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  44. Room Impulse Responses help attackers to evade Deep Fake Detection
    2024/09/23 by Hieu-Thi Luong, Luong, Hieu-Thi, Duc-Tuan Truong +5 · 2 citations
    Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  45. MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
    2025/11/02 by Yayue Deng, Guoqiang Hu, Deng, Yayue +15 · 3 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  46. The ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge (ICSRC): Dataset, Tracks, Baseline and Results
    2022/11/03 by Ao Zhang, Fan Yu, Zhang, Ao +17 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and dialogue systems
  47. Noise robust distillation of self-supervised speech models via correlation metrics
    2023/12/19 by Fabian Ritter-Gutierrez, Kuan-Po Huang, Ritter-Gutierrez, Fabian +11 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  48. Internal Language Model Estimation based Language Model Fusion for Cross-Domain Code-Switching Speech Recognition
    2022/07/09 by Yizhou Peng, Peng, Yizhou, Yufei Liu +11 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  49. Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models
    2025/10/10 by Donghang Wu, Wu, Donghang, Haoyang Zhang +20 · 4 citations
    Computer Science · Psychology · #Action Observation and Synchronization #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
  50. Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
    2025/09/29 by Hexin Liu, Haoyang Zhang, Liu, Hexin +11 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing