vix.ing · top · new · best · stats · spec

Chng, Eng Siong

  1. SpEx+: A Complete Time Domain Speaker Extraction Network
    2020/05/10 by Meng Ge, Chenglin Xu, Ge, Meng +9 · 16 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  2. Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
    2025/01/13 by Ma, Ziyang, Chen, Zhuo, Wang, Yuping +2 · 32 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  3. HyPoradise: An Open Baseline for Generative Speech Recognition with Large Language Models
    2023/09/27 by Chen, Chen, Hu, Yuchen, Yang, Chao-Han Huck +3 · 11 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  4. Speech Emotion Recognition with Co-Attention based Multi-level Acoustic Information
    2022/03/29 by Heqing Zou, Yuke Si, Zou, Heqing +7 · 7 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
    2024/06/02 by Chen Chen, Chen, Chen, Yuchen Hu +9 · 13 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Robotics and Automated Systems #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  6. L-SpEx: Localized Target Speaker Extraction
    2022/02/21 by Meng Ge, Chenglin Xu, Ge, Meng +9 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. Noise-robust Speech Recognition with 10 Minutes Unparalleled In-domain Data
    2022/03/29 by Chen Chen, Nana Hou, Chen, Chen +7 · 5 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  8. Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
    2024/07/02 by Yuchen Hu, Chen Chen, Hu, Yuchen +7 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  9. Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
    2024/05/16 by Yuchen Hu, Hu, Yuchen, Chen Chen +9 · 7 citations
    Computer Science · #Speech Recognition and Synthesis
  10. Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
    2025/01/27 by Chen Chen, Chen, Chen, Yuchen Hu +13 · 13 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  11. GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
    2024/02/10 by Hu, Yuchen, Chen, Chen, Yang, Chao-Han Huck +4 · 6 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  12. Adapting OpenAI's Whisper for Speech Recognition on Code-Switch Mandarin-English SEAME and ASRU2019 Datasets
    2023/11/29 by Yuhang Yang, Yizhou Peng, Yang, Yuhang +7 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  13. On the End-to-End Solution to Mandarin-English Code-switching Speech Recognition
    2018/11/01 by Zeng, Zhiping, Khassanov, Yerbolat, Pham, Van Tung +3 · 3 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  14. Cross-Modality and Within-Modality Regularization for Audio-Visual DeepFake Detection
    2024/01/11 by Zou, Heqing, Shen, Meng, Hu, Yuchen +3 · 6 citations
    #FOS: Computer and information sciences #Multimedia (cs.MM)
  15. Leveraging Modality-specific Representations for Audio-visual Speech Recognition via Reinforcement Learning
    2022/12/10 by Chen, Chen, Hu, Yuchen, Zhang, Qiang +3 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  16. LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
    2024/09/23 by Hieu-Thi Luong, Luong, Hieu-Thi, Haoyang Li +7 · 9 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  17. Multi-stage Speaker Extraction with Utterance and Frame-Level Reference Signals
    2020/11/19 by Meng Ge, Chenglin Xu, Ge, Meng +9 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  18. Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition
    2023/06/18 by Hu, Yuchen, Li, Ruizhe, Chen, Chen +3 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  19. Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition
    2023/05/16 by Yu‐Chen Hu, Ruizhe Li, Hu, Yuchen +9 · 4 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  20. Towards Audio Codec-based Speech Separation
    2024/06/18 by Jia Qi Yip, Yip, Jia Qi, Shengkui Zhao +7 · 5 citations
    Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  21. Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model
    2024/02/16 by Xiangyu Zhang, Daijiao Liu, Zhang, Xiangyu +13 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech and Audio Processing #electronic engineering #information engineering
  22. I2CR: Improving Noise Robustness on Keyword Spotting Using Inter-Intra Contrastive Regularization
    2022/09/14 by Dianwen Ng, Ng, Dianwen, Jia Qi Yip +15 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  23. deHuBERT: Disentangling Noise in a Self-supervised Model for Robust Speech Recognition
    2023/02/28 by Ng, Dianwen, Zhang, Ruixi, Yip, Jia Qi +7 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  24. Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models
    2024/05/23 by Yuchen Hu, Chen Chen, Hu, Yuchen +11 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  25. Minimum word error training for non-autoregressive Transformer-based code-switching ASR
    2021/10/07 by Yizhou Peng, Peng, Yizhou, Jicheng Zhang +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Wireless Signal Modulation Classification #electronic engineering #information engineering
  26. Aligning Speech to Languages to Enhance Code-switching Speech Recognition
    2024/03/09 by Liu, Hexin, Zhang, Xiangyu, Zhang, Haoyang +4 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  27. Speech Separation using Neural Audio Codecs with Embedding Loss
    2024/11/27 by Jia Qi Yip, Yip, Jia Qi, Chin Yuen Kwok +5 · 4 citations
    Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  28. FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
    2025/07/25 by Y Peng, Peng, Yizhou, Yi Chao +11 · 5 citations
    Computer Science · Decision Sciences · #Speech and dialogue systems #Personal Information Management and User Behavior #Topic Modeling
  29. Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
    2024/09/15 by Chao-Han Huck Yang, Taejin Park, Yang, Chao-Han Huck +39 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  30. Adapting BERT for Word Sense Disambiguation with Gloss Selection\n Objective and Example Sentences
    2020/09/24 by Boon Peng Yap, Yap, Boon Peng, Andrew Y. Koh +3 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  31. Multitask-Based Joint Learning Approach To Robust ASR For Radio Communication Speech
    2021/07/22 by Ma, Duo, Hou, Nana, Pham, Van Tung +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  32. Automated Audio Captioning using Transfer Learning and Reconstruction Latent Space Similarity Regularization
    2021/08/10 by Koh, Andrew, Xue, Fuzhao, Chng, Eng Siong · 1 citation
    #68T50 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Sound (cs.SD) #electronic engineering #information engineering
  33. Interactive Feature Fusion for End-to-End Noise-Robust Speech Recognition
    2021/10/11 by Hu, Yuchen, Hou, Nana, Chen, Chen +1 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  34. Distilling a speech and music encoder with task arithmetic
    2025/05/19 by Fabian Ritter-Gutierrez, Ritter-Gutierrez, Fabian, Yi‐Cheng Lin +10 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  35. Interactive Audio-text Representation for Automated Audio Captioning with Contrastive Learning
    2022/03/29 by Chen Chen, Chen, Chen, Nana Hou +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  36. Dataset-Distillation Generative Model for Speech Emotion Recognition
    2024/06/05 by Fabian Ritter-Gutierrez, Kuan-Po Huang, Ritter-Gutierrez, Fabian +11 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Advanced Data Compression Techniques
  37. Dual-Path Style Learning for End-to-End Noise-Robust Speech Recognition
    2022/03/28 by Hu, Yuchen, Hou, Nana, Chen, Chen +1 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  38. Self-critical Sequence Training for Automatic Speech Recognition
    2022/04/13 by Chen, Chen, Hu, Yuchen, Hou, Nana +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  39. Intermediate-layer output Regularization for Attention-based Speech Recognition with Shared Decoder
    2022/07/09 by Jicheng Zhang, Yizhou Peng, Zhang, Jicheng +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  40. A Vocoder-free WaveNet Voice Conversion with Non-Parallel Data
    2019/02/11 by Xiaohai Tian, Tian, Xiaohai, Eng Siong Chng +3 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  41. ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge
    2024/01/07 by Wang, He, Guo, Pengcheng, Li, Yue +13 · 2 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  42. Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech Separation
    2023/02/22 by Yu‐Chen Hu, Hu, Yuchen, Chen Chen +7 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  43. Leveraging Audio-Tagging Assisted Sound Event Detection using Weakified Strong Labels and Frequency Dynamic Convolutions
    2023/04/25 by Tanmay Khandelwal, Rohan Kumar Das, Khandelwal, Tanmay +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Speech and Audio Processing #electronic engineering #information engineering
  44. UniS-MMC: Multimodal Classification via Unimodality-supervised Multimodal Contrastive Learning
    2023/05/16 by Heqing Zou, Zou, Heqing, Meng Shen +9 · 1 citation
    Computer Science · Engineering · #Advanced Chemical Sensor Technologies #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Text and Document Classification Technologies
  45. Rainbow Keywords: Efficient Incremental Learning for Online Spoken Keyword Spotting
    2022/03/30 by Xiao Yang, Nana Hou, Xiao, Yang +3 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  46. Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding
    2025/05/12 by Dianwen Ng, Ng, Dianwen, Kun Zhou +9 · 3 citations
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  47. SPGM: Prioritizing Local Features for enhanced speech separation performance
    2023/09/22 by Jia Qi Yip, Shengkui Zhao, Yip, Jia Qi +19 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  48. Continual Learning For On-Device Environmental Sound Classification
    2022/07/15 by Xiao, Yang, Liu, Xubo, King, James +4 · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  49. Are Soft Prompts Good Zero-shot Learners for Speech Recognition?
    2023/09/18 by Dianwen Ng, Ng, Dianwen, Chong Zhang +17 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  50. Generative error correction for code-switching speech recognition using large language models
    2023/10/17 by Chen Chen, Chen, Chen, Yuchen Hu +9 · 1 citation
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems
  51. MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
    2025/11/02 by Deng, Yayue, Hu, Guoqiang, Sun, Haiyang +6 · 3 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  52. The ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge (ICSRC): Dataset, Tracks, Baseline and Results
    2022/11/03 by Ao Zhang, Zhang, Ao, Fan Yu +17 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and dialogue systems
  53. Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English
    2025/05/20 by Zhang, Haoyang, Liu, Hexin, Zhang, Xiangyu +7 · 1 citation
    #68T10 #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Sound (cs.SD) #electronic engineering #information engineering
  54. Internal Language Model Estimation based Language Model Fusion for Cross-Domain Code-Switching Speech Recognition
    2022/07/09 by Yizhou Peng, Peng, Yizhou, Yufei Liu +11 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  55. Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models
    2025/10/10 by Donghang Wu, Haoyang Zhang, Wu, Donghang +20 · 4 citations
    Computer Science · Psychology · #Action Observation and Synchronization #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
  56. Speechless: Speech Instruction Training Without Speech for Low Resource Languages
    2025/05/23 by Dao, Alan, Vu, Dinh Bach, Ha, Huy Hoang +6 · 3 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  57. Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
    2025/10/14 by Ma, Ziyang, Xu, Ruiyang, Xing, Zhenghao +9 · 2 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM) #Sound (cs.SD)
  58. Independent language modeling architecture for end-to-end ASR
    2019/11/25 by Van Tung Pham, Haihua Xu, Pham, Van Tung +13 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  59. Step-Audio-R1 Technical Report
    2025/11/19 by Tian, Fei, Zhang, Xiangyu Tony, Zhang, Yuxin +14 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #H.5.5 #I.2.6 #I.2.7 #Multimodal Machine Learning Applications #Sound (cs.SD) #Speech and dialogue systems
  60. Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
    2025/09/29 by Hexin Liu, Liu, Hexin, Haoyang Zhang +11 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing
  61. Noise robust distillation of self-supervised speech models via correlation metrics
    2023/12/19 by Fabian Ritter-Gutierrez, Ritter-Gutierrez, Fabian, Kuan-Po Huang +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering