vix.ing · top · new · best · stats · spec

Li, Jinyu

  1. Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
    2023/01/05 by Wang, Chengyi, Chen, Sanyuan, Wu, Yu +10 · 251 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  2. NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
    2024/03/05 by Zeqian Ju, Ju, Zeqian, Yuancheng Wang +35 · 118 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
    2025/03/03 by Microsoft, :, Abdelrahman Abouelenin +147 · 226 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
  4. SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
    2021/10/14 by Junyi Ao, Ao, Junyi, Rui Wang +25 · 35 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
    2024/06/08 by Chen, Sanyuan, Liu, Shujie, Zhou, Long +6 · 62 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  6. Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
    2023/03/07 by Ziqiang Zhang, Long Zhou, Zhang, Ziqiang +23 · 40 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  7. WavLLM: Towards Robust and Adaptive Speech Large Language Model
    2024/03/31 by Shujie Hu, Hu, Shujie, Long Zhou +20 · 53 citations
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  8. On decoder-only architecture for speech-to-text and large language model integration
    2023/07/08 by Jian Wu, Yashesh Gaur, Wu, Jian +19 · 38 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  9. Recent Advances in End-to-End Automatic Speech Recognition
    2021/11/02 by Jinyu Li, Li, Jinyu · 28 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. Continuous speech separation: dataset and analysis
    2020/01/30 by Zhuo Chen, Takuya Yoshioka, Chen, Zhuo +15 · 22 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
  11. UniSpeech-SAT: Universal Speech Representation Learning with Speaker Aware Pre-Training
    2021/10/12 by Sanyuan Chen, Yu Wu, Chen, Sanyuan +18 · 19 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. Autoregressive Speech Synthesis without Vector Quantization
    2024/07/11 by Lingwei Meng, Meng, Lingwei, Long Zhou +21 · 36 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
  13. SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
    2023/08/14 by Xiaofei Wang, Wang, Xiaofei, Manthan Thakker +17 · 23 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Topic Modeling
  14. Integration of speech separation, diarization, and recognition for\n multi-speaker meetings: System description, comparison, and analysis
    2020/11/03 by Desh Raj, Raj, Desh, Pavel Denisov +25 · 14 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  15. VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
    2024/06/12 by Bing Han, Han, Bing, Long Zhou +17 · 23 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  16. Streaming Multi-Talker ASR with Token-Level Serialized Output Training
    2022/02/02 by Kanda, Naoyuki, Wu, Jian, Wu, Yu +7 · 10 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  17. SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
    2024/12/20 by Wenxi Chen, Ziyang Ma, Chen, Wenxi +27 · 31 citations
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
  18. Prompting Large Language Models for Zero-Shot Domain Adaptation in Speech Recognition
    2023/06/28 by Yuang Li, Li, Yuang, Yu Wu +5 · 12 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Signal Processing (eess.SP) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  19. PillarNeXt: Rethinking Network Designs for 3D Object Detection in LiDAR Point Clouds
    2023/05/08 by Jinyu Li, Li, Jinyu, Chenxu Luo +3 · 11 citations
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Industrial Vision Systems and Defect Detection
  20. Developing Real-time Streaming Transformer Transducer for Speech Recognition on Large-scale Dataset
    2020/10/22 by Xie Chen, Yu Wu, Chen, Xie +7 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  21. ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation
    2024/10/27 by Zongyi Li, Shujie Hu, Li, Zongyi +16 · 18 citations
    Computer Science · #Generative Adversarial Networks and Image Synthesis #Advanced Image Processing Techniques #Image and Video Quality Assessment
  22. Why does Self-Supervised Learning for Speech Recognition Benefit Speaker Recognition?
    2022/04/27 by Chen, Sanyuan, Wu, Yu, Wang, Chengyi +8 · 8 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  23. Continuous Speech Separation with Conformer
    2020/08/13 by Sanyuan Chen, Chen, Sanyuan, Yu Wu +15 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  24. VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
    2023/05/25 by Wang, Tianrui, Zhou, Long, Zhang, Ziqiang +6 · 9 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  25. Streaming Speaker-Attributed ASR with Token-Level Speaker Embeddings
    2022/03/30 by Naoyuki Kanda, Kanda, Naoyuki, Wu, Jian +16 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  26. Wav2vec-Switch: Contrastive Learning from Original-noisy Speech Pairs for Robust Speech Recognition
    2021/10/11 by Yiming Wang, Wang, Yiming, Jinyu Li +9 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  27. TS3-Codec: Transformer-Based Simple Streaming Single Codec
    2024/11/27 by Haibin Wu, Wu, Haibin, Naoyuki Kanda +5 · 15 citations
    Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #Digital Filter Design and Implementation #FOS: Electrical engineering #electronic engineering #information engineering
  28. Don't shoot butterfly with rifles: Multi-channel Continuous Speech Separation with Early Exit Transformer
    2020/10/23 by Sanyuan Chen, Chen, Sanyuan, Yu Wu +9 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  29. Large-Scale Streaming End-to-End Speech Translation with Neural Transducers
    2022/04/11 by Xue, Jian, Wang, Peidong, Li, Jinyu +2 · 7 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  30. CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations
    2024/04/10 by Leying Zhang, Yao Qian, Zhang, Leying +21 · 10 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  31. SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data
    2022/09/30 by Zhang, Ziqiang, Chen, Sanyuan, Zhou, Long +8 · 6 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  32. Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
    2024/07/17 by Haibin Wu, Wu, Haibin, Xiaofei Wang +19 · 11 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Signal Processing (eess.SP) #Speech Recognition and Synthesis #electronic engineering #information engineering
  33. On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition
    2020/05/28 by Jinyu Li, Yu Wu, Li, Jinyu +9 · 6 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
  34. Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning
    2025/02/19 by Zhao, Rui, Yuan, Qirui, Li, Jinyu +4 · 17 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  35. RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
    2024/04/04 by Detai Xin, Xin, Detai, Xu Tan +19 · 9 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  36. Accurate and Structured Pruning for Efficient Automatic Speech Recognition
    2023/05/31 by Huiqiang Jiang, Li Lyna Zhang, Jiang, Huiqiang +17 · 6 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
  37. Boosting Large Language Model for Speech Synthesis: An Empirical Study
    2023/12/30 by Hongkun Hao, Long Zhou, Hao, Hongkun +11 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  38. Accelerating Transducers through Adjacent Token Merging
    2023/06/28 by Li, Yuang, Wu, Yu, Li, Jinyu +1 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Signal Processing (eess.SP) #electronic engineering #information engineering
  39. Exploring Transformers for Large-Scale Speech Recognition
    2020/05/19 by Lu, Liang, Liu, Changliang, Li, Jinyu +1 · 3 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  40. COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning
    2023/11/03 by Jing Pan, Pan, Jing, Jian Wu +11 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  41. Semantic Mask for Transformer based End-to-End Speech Recognition
    2019/12/06 by Chengyi Wang, Wang, Chengyi, Yu Wu +17 · 9 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  42. Factorized Neural Transducer for Efficient Language Model Adaptation
    2021/09/27 by Xie Chen, Chen, Xie, Zhong Meng +5 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Music and Audio Processing
  43. Discrete Audio Tokens: More Than a Survey!
    2025/06/12 by Pooneh Mousavi, Gallil Maimon, Mousavi, Pooneh +39 · 14 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  44. Low Latency End-to-End Streaming Speech Recognition with a Scout Network
    2020/03/23 by Chengyi Wang, Yu Wu, Wang, Chengyi +11 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  45. Supervision-Guided Codebooks for Masked Prediction in Speech Pre-training
    2022/06/21 by Chengyi Wang, Yiming Wang, Wang, Chengyi +11 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  46. Simulating realistic speech overlaps improves multi-talker ASR
    2022/10/27 by Muqiao Yang, Yang, Muqiao, Naoyuki Kanda +12 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  47. TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
    2024/05/28 by Chenyang Le, Yao Qian, Le, Chenyang +19 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  48. Speech separation with large-scale self-supervised learning
    2022/11/09 by Zhuo Chen, Naoyuki Kanda, Chen, Zhuo +14 · 3 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  49. LongFNT: Long-form Speech Recognition with Factorized Neural Transducer
    2022/11/17 by Xun Gong, Gong, Xun, Yu Wu +11 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  50. A Configurable Multilingual Model is All You Need to Recognize All Languages
    2021/07/13 by Long Zhou, Zhou, Long, Jinyu Li +5 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  51. A Light-weight contextual spelling correction model for customizing transducer-based speech recognition systems
    2021/08/17 by Xiaoqiang Wang, Yanqing Liu, Wang, Xiaoqiang +5 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  52. Self-Supervised Learning for speech recognition with Intermediate layer supervision
    2021/12/16 by Wang, Chengyi, Wu, Yu, Chen, Sanyuan +4 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  53. Improving Noise Robustness of Contrastive Speech Representation Learning with Speech Reconstruction
    2021/10/28 by Heming Wang, Yao Qian, Wang, Heming +15 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  54. Investigation of Practical Aspects of Single Channel Speech Separation for ASR
    2021/07/05 by Jian Wu, Wu, Jian, Zhuo Chen +13 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  55. Continuous Streaming Multi-Talker ASR with Dual-path Transducers
    2021/09/17 by Desh Raj, Liang Lu, Raj, Desh +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  56. FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
    2025/02/16 by Wang, Hui, Liu, Shujie, Meng, Lingwei +9 · 10 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  57. ResidualTransformer: Residual Low-Rank Learning with Weight-Sharing for Transformer Layers
    2023/10/03 by Wang, Yiming, Li, Jinyu · 3 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  58. Towards Contextual Spelling Correction for Customization of End-to-end Speech Recognition Systems
    2022/03/02 by Xiaoqiang Wang, Wang, Xiaoqiang, Yanqing Liu +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  59. D3FlowSLAM: Self-Supervised Dynamic SLAM with Flow Motion Decomposition and DINO Guidance
    2022/07/18 by Xingyuan Yu, Weicai Ye, Yu, Xingyuan +12 · 2 citations
    Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics (cs.RO) #Robotics and Sensor-Based Localization
  60. Fast and accurate factorized neural transducer for text adaption of end-to-end speech recognition models
    2022/12/05 by Rui Zhao, Jian Xue, Zhao, Rui +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  61. Developing RNN-T Models Surpassing High-Performance Hybrid Models with Customization Capability
    2020/07/30 by Li, Jinyu, Zhao, Rui, Meng, Zhong +8 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  62. Improving Contextual Spelling Correction by External Acoustics Attention and Semantic Aware Data Augmentation
    2023/02/22 by Xiaoqiang Wang, Yanqing Liu, Wang, Xiaoqiang +5 · 2 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  63. End-to-End Attention based Text-Dependent Speaker Verification
    2017/01/03 by Shi-Xiong Zhang, Zhuo Chen, Zhang, Shi-Xiong +7 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (stat.ML) #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
  64. AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM
    2024/12/02 by Ruchao Fan, Bo Ren, Fan, Ruchao +9 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  65. Large-Scale Domain Adaptation via Teacher-Student Learning
    2017/08/17 by Li, Jinyu, Seltzer, Michael L., Wang, Xi +2 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  66. Improved training for online end-to-end speech recognition systems
    2017/11/06 by Suyoun Kim, Michael L. Seltzer, Kim, Suyoun +5 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
  67. V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
    2024/11/29 by Jeongsoo Choi, Jihoon Kim, Choi, Jeongsoo +7 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  68. VarArray Meets t-SOT: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition
    2022/09/12 by Naoyuki Kanda, Kanda, Naoyuki, Xiaofei Wang +8 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  69. Speaker Adaptation for End-to-End CTC Models
    2019/01/04 by Ke Li, Jinyu Li, Li, Ke +7 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
  70. A Weakly-Supervised Streaming Multilingual Speech Model with Truly Zero-Shot Capability
    2022/11/04 by Jian Xue, Xue, Jian, Peidong Wang +5 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  71. LAMASSU: Streaming Language-Agnostic Multilingual Speech Recognition and Translation Using Neural Transducers
    2022/11/05 by Peidong Wang, Eric Sun, Wang, Peidong +13 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  72. Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
    2024/09/06 by Jiaqi Li, Dongmei Wang, Li, Jiaqi +29 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing
  73. Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
    2024/10/17 by Sreyan Ghosh, Mohammad Sadegh Rasooli, Ghosh, Sreyan +11 · 4 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Fault Detection and Control Systems #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  74. An End-to-end Architecture of Online Multi-channel Speech Separation
    2020/09/07 by Jian Wu, Wu, Jian, Zhuo Chen +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  75. Feature Learning in Deep Neural Networks - Studies on Speech Recognition\n Tasks
    2013/01/16 by Dong Yu, Yu, Dong, Michael L. Seltzer +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Neural and Evolutionary Computing (cs.NE) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  76. Streaming Multi-talker Speech Recognition with Joint Speaker Identification
    2021/04/05 by Lu, Liang, Kanda, Naoyuki, Li, Jinyu +1 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Sound (cs.SD)
  77. On Addressing Practical Challenges for RNN-Transducer
    2021/04/27 by Zhao, Rui, Xue, Jian, Li, Jinyu +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  78. Continuous Speech Separation with Recurrent Selective Attention Network
    2021/10/28 by Yixuan Zhang, Zhang, Yixuan, Zhuo Chen +10 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  79. Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Speech Data
    2022/03/31 by Junyi Ao, Ao, Junyi, Ziqiang Zhang +17 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  80. Endpoint Detection for Streaming End-to-End Multi-talker ASR
    2022/01/24 by Lu, Liang, Li, Jinyu, Gong, Yifan · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  81. Total-Duration-Aware Duration Modeling for Text-to-Speech Systems
    2024/06/06 by Şefik Emre Eskimez, Eskimez, Sefik Emre, Xiaofei Wang +21 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering
  82. SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training
    2022/10/07 by Ziqiang Zhang, Long Zhou, Zhang, Ziqiang +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  83. Self-supervised learning with bi-label masked speech prediction for streaming multi-talker speech recognition
    2022/11/10 by Zili Huang, Huang, Zili, Zhuo Chen +14 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  84. CTCBERT: Advancing Hidden-unit BERT with CTC Objectives
    2022/10/16 by Ruchao Fan, Yiming Wang, Fan, Ruchao +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  85. Streaming, fast and accurate on-device Inverse Text Normalization for Automatic Speech Recognition
    2022/11/07 by Yashesh Gaur, Nick Kibre, Gaur, Yashesh +13 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems
  86. Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
    2025/04/14 by Yifan Yang, Yang, Yifan, Shujie Liu +22 · 7 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  87. CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
    2025/06/01 by Yao Qian, Zhang, Leying, Xiaofei Wang +17 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  88. Building High-accuracy Multilingual ASR with Gated Language Experts and Curriculum Training
    2023/03/01 by Eric Sun, Sun, Eric, Jinyu Li +19 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  89. t-SOT FNT: Streaming Multi-talker ASR with Text-only Domain Adaptation Capability
    2023/09/15 by Wu, Jian, Kanda, Naoyuki, Yoshioka, Takuya +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  90. High-Accuracy and Low-Latency Speech Recognition with Two-Head Contextual Layer Trajectory LSTM Model
    2020/03/17 by Jinyu Li, Li, Jinyu, Rui Zhao +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  91. An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
    2024/06/09 by Xiaofei Wang, Wang, Xiaofei, Şefik Emre Eskimez +19 · 1 citation
    Engineering · Computer Science · #Advanced Adaptive Filtering Techniques #Speech and Audio Processing #Blind Source Separation Techniques
  92. Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
    2024/12/20 by Y. F. Yang, Shujie Liu, Yang, Yifan +22 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering
  93. StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
    2025/06/14 by Hui Wang, Wang, Hui, Shujie Liu +16 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  94. Addressing speaker gender bias in large scale speech translation systems
    2025/01/10 by Shubham Bansal, Vikas Joshi, Bansal, Shubham +7 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems