Li, Jinyu
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
2023/01/05 by Wang, Chengyi, Chen, Sanyuan, Wu, Yu +10 · 251 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
2024/03/05 by Zeqian Ju, Ju, Zeqian, Yuancheng Wang +35 · 118 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
2025/03/03 by Microsoft, :, Abdelrahman Abouelenin +147 · 226 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
- SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
2021/10/14 by Junyi Ao, Ao, Junyi, Rui Wang +25 · 35 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
2024/06/08 by Chen, Sanyuan, Liu, Shujie, Zhou, Long +6 · 62 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
2023/03/07 by Ziqiang Zhang, Long Zhou, Zhang, Ziqiang +23 · 40 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- WavLLM: Towards Robust and Adaptive Speech Large Language Model
2024/03/31 by Shujie Hu, Hu, Shujie, Long Zhou +20 · 53 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- On decoder-only architecture for speech-to-text and large language model integration
2023/07/08 by Jian Wu, Yashesh Gaur, Wu, Jian +19 · 38 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Recent Advances in End-to-End Automatic Speech Recognition
2021/11/02 by Jinyu Li, Li, Jinyu · 28 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Continuous speech separation: dataset and analysis
2020/01/30 by Zhuo Chen, Takuya Yoshioka, Chen, Zhuo +15 · 22 citations
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
- UniSpeech-SAT: Universal Speech Representation Learning with Speaker Aware Pre-Training
2021/10/12 by Sanyuan Chen, Yu Wu, Chen, Sanyuan +18 · 19 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Autoregressive Speech Synthesis without Vector Quantization
2024/07/11 by Lingwei Meng, Meng, Lingwei, Long Zhou +21 · 36 citations
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
- SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
2023/08/14 by Xiaofei Wang, Wang, Xiaofei, Manthan Thakker +17 · 23 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Topic Modeling
- Integration of speech separation, diarization, and recognition for\n multi-speaker meetings: System description, comparison, and analysis
2020/11/03 by Desh Raj, Raj, Desh, Pavel Denisov +25 · 14 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
2024/06/12 by Bing Han, Han, Bing, Long Zhou +17 · 23 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Streaming Multi-Talker ASR with Token-Level Serialized Output Training
2022/02/02 by Kanda, Naoyuki, Wu, Jian, Wu, Yu +7 · 10 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
2024/12/20 by Wenxi Chen, Ziyang Ma, Chen, Wenxi +27 · 31 citations
Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
- Prompting Large Language Models for Zero-Shot Domain Adaptation in Speech Recognition
2023/06/28 by Yuang Li, Li, Yuang, Yu Wu +5 · 12 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Signal Processing (eess.SP) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- PillarNeXt: Rethinking Network Designs for 3D Object Detection in LiDAR Point Clouds
2023/05/08 by Jinyu Li, Li, Jinyu, Chenxu Luo +3 · 11 citations
Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Industrial Vision Systems and Defect Detection
- Developing Real-time Streaming Transformer Transducer for Speech Recognition on Large-scale Dataset
2020/10/22 by Xie Chen, Yu Wu, Chen, Xie +7 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation
2024/10/27 by Zongyi Li, Shujie Hu, Li, Zongyi +16 · 18 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis #Advanced Image Processing Techniques #Image and Video Quality Assessment
- Why does Self-Supervised Learning for Speech Recognition Benefit Speaker Recognition?
2022/04/27 by Chen, Sanyuan, Wu, Yu, Wang, Chengyi +8 · 8 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Continuous Speech Separation with Conformer
2020/08/13 by Sanyuan Chen, Chen, Sanyuan, Yu Wu +15 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
2023/05/25 by Wang, Tianrui, Zhou, Long, Zhang, Ziqiang +6 · 9 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Streaming Speaker-Attributed ASR with Token-Level Speaker Embeddings
2022/03/30 by Naoyuki Kanda, Kanda, Naoyuki, Wu, Jian +16 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Wav2vec-Switch: Contrastive Learning from Original-noisy Speech Pairs for Robust Speech Recognition
2021/10/11 by Yiming Wang, Wang, Yiming, Jinyu Li +9 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- TS3-Codec: Transformer-Based Simple Streaming Single Codec
2024/11/27 by Haibin Wu, Wu, Haibin, Naoyuki Kanda +5 · 15 citations
Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #Digital Filter Design and Implementation #FOS: Electrical engineering #electronic engineering #information engineering
- Don't shoot butterfly with rifles: Multi-channel Continuous Speech Separation with Early Exit Transformer
2020/10/23 by Sanyuan Chen, Chen, Sanyuan, Yu Wu +9 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Large-Scale Streaming End-to-End Speech Translation with Neural Transducers
2022/04/11 by Xue, Jian, Wang, Peidong, Li, Jinyu +2 · 7 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations
2024/04/10 by Leying Zhang, Yao Qian, Zhang, Leying +21 · 10 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data
2022/09/30 by Zhang, Ziqiang, Chen, Sanyuan, Zhou, Long +8 · 6 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
2024/07/17 by Haibin Wu, Wu, Haibin, Xiaofei Wang +19 · 11 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Signal Processing (eess.SP) #Speech Recognition and Synthesis #electronic engineering #information engineering
- On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition
2020/05/28 by Jinyu Li, Yu Wu, Li, Jinyu +9 · 6 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
- Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning
2025/02/19 by Zhao, Rui, Yuan, Qirui, Li, Jinyu +4 · 17 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
2024/04/04 by Detai Xin, Xin, Detai, Xu Tan +19 · 9 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Accurate and Structured Pruning for Efficient Automatic Speech Recognition
2023/05/31 by Huiqiang Jiang, Li Lyna Zhang, Jiang, Huiqiang +17 · 6 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
- Boosting Large Language Model for Speech Synthesis: An Empirical Study
2023/12/30 by Hongkun Hao, Long Zhou, Hao, Hongkun +11 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Accelerating Transducers through Adjacent Token Merging
2023/06/28 by Li, Yuang, Wu, Yu, Li, Jinyu +1 · 4 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Signal Processing (eess.SP) #electronic engineering #information engineering
- Exploring Transformers for Large-Scale Speech Recognition
2020/05/19 by Lu, Liang, Liu, Changliang, Li, Jinyu +1 · 3 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning
2023/11/03 by Jing Pan, Pan, Jing, Jian Wu +11 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Semantic Mask for Transformer based End-to-End Speech Recognition
2019/12/06 by Chengyi Wang, Wang, Chengyi, Yu Wu +17 · 9 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- Factorized Neural Transducer for Efficient Language Model Adaptation
2021/09/27 by Xie Chen, Chen, Xie, Zhong Meng +5 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Music and Audio Processing
- Discrete Audio Tokens: More Than a Survey!
2025/06/12 by Pooneh Mousavi, Gallil Maimon, Mousavi, Pooneh +39 · 14 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Low Latency End-to-End Streaming Speech Recognition with a Scout Network
2020/03/23 by Chengyi Wang, Yu Wu, Wang, Chengyi +11 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Supervision-Guided Codebooks for Masked Prediction in Speech Pre-training
2022/06/21 by Chengyi Wang, Yiming Wang, Wang, Chengyi +11 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Simulating realistic speech overlaps improves multi-talker ASR
2022/10/27 by Muqiao Yang, Yang, Muqiao, Naoyuki Kanda +12 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
2024/05/28 by Chenyang Le, Yao Qian, Le, Chenyang +19 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Speech separation with large-scale self-supervised learning
2022/11/09 by Zhuo Chen, Naoyuki Kanda, Chen, Zhuo +14 · 3 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- LongFNT: Long-form Speech Recognition with Factorized Neural Transducer
2022/11/17 by Xun Gong, Gong, Xun, Yu Wu +11 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- A Configurable Multilingual Model is All You Need to Recognize All Languages
2021/07/13 by Long Zhou, Zhou, Long, Jinyu Li +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- A Light-weight contextual spelling correction model for customizing transducer-based speech recognition systems
2021/08/17 by Xiaoqiang Wang, Yanqing Liu, Wang, Xiaoqiang +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Self-Supervised Learning for speech recognition with Intermediate layer supervision
2021/12/16 by Wang, Chengyi, Wu, Yu, Chen, Sanyuan +4 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Improving Noise Robustness of Contrastive Speech Representation Learning with Speech Reconstruction
2021/10/28 by Heming Wang, Yao Qian, Wang, Heming +15 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Investigation of Practical Aspects of Single Channel Speech Separation for ASR
2021/07/05 by Jian Wu, Wu, Jian, Zhuo Chen +13 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Continuous Streaming Multi-Talker ASR with Dual-path Transducers
2021/09/17 by Desh Raj, Liang Lu, Raj, Desh +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
2025/02/16 by Wang, Hui, Liu, Shujie, Meng, Lingwei +9 · 10 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- ResidualTransformer: Residual Low-Rank Learning with Weight-Sharing for Transformer Layers
2023/10/03 by Wang, Yiming, Li, Jinyu · 3 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Towards Contextual Spelling Correction for Customization of End-to-end Speech Recognition Systems
2022/03/02 by Xiaoqiang Wang, Wang, Xiaoqiang, Yanqing Liu +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- D3FlowSLAM: Self-Supervised Dynamic SLAM with Flow Motion Decomposition and DINO Guidance
2022/07/18 by Xingyuan Yu, Weicai Ye, Yu, Xingyuan +12 · 2 citations
Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics (cs.RO) #Robotics and Sensor-Based Localization
- Fast and accurate factorized neural transducer for text adaption of end-to-end speech recognition models
2022/12/05 by Rui Zhao, Jian Xue, Zhao, Rui +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Developing RNN-T Models Surpassing High-Performance Hybrid Models with Customization Capability
2020/07/30 by Li, Jinyu, Zhao, Rui, Meng, Zhong +8 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Improving Contextual Spelling Correction by External Acoustics Attention and Semantic Aware Data Augmentation
2023/02/22 by Xiaoqiang Wang, Yanqing Liu, Wang, Xiaoqiang +5 · 2 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- End-to-End Attention based Text-Dependent Speaker Verification
2017/01/03 by Shi-Xiong Zhang, Zhuo Chen, Zhang, Shi-Xiong +7 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (stat.ML) #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
- AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM
2024/12/02 by Ruchao Fan, Bo Ren, Fan, Ruchao +9 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Large-Scale Domain Adaptation via Teacher-Student Learning
2017/08/17 by Li, Jinyu, Seltzer, Michael L., Wang, Xi +2 · 1 citation
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- Improved training for online end-to-end speech recognition systems
2017/11/06 by Suyoun Kim, Michael L. Seltzer, Kim, Suyoun +5 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
- V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
2024/11/29 by Jeongsoo Choi, Jihoon Kim, Choi, Jeongsoo +7 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- VarArray Meets t-SOT: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition
2022/09/12 by Naoyuki Kanda, Kanda, Naoyuki, Xiaofei Wang +8 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Speaker Adaptation for End-to-End CTC Models
2019/01/04 by Ke Li, Jinyu Li, Li, Ke +7 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
- A Weakly-Supervised Streaming Multilingual Speech Model with Truly Zero-Shot Capability
2022/11/04 by Jian Xue, Xue, Jian, Peidong Wang +5 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- LAMASSU: Streaming Language-Agnostic Multilingual Speech Recognition and Translation Using Neural Transducers
2022/11/05 by Peidong Wang, Eric Sun, Wang, Peidong +13 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
- Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
2024/09/06 by Jiaqi Li, Dongmei Wang, Li, Jiaqi +29 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing
- Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
2024/10/17 by Sreyan Ghosh, Mohammad Sadegh Rasooli, Ghosh, Sreyan +11 · 4 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Fault Detection and Control Systems #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- An End-to-end Architecture of Online Multi-channel Speech Separation
2020/09/07 by Jian Wu, Wu, Jian, Zhuo Chen +13 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Feature Learning in Deep Neural Networks - Studies on Speech Recognition\n Tasks
2013/01/16 by Dong Yu, Yu, Dong, Michael L. Seltzer +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Neural and Evolutionary Computing (cs.NE) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Streaming Multi-talker Speech Recognition with Joint Speaker Identification
2021/04/05 by Lu, Liang, Kanda, Naoyuki, Li, Jinyu +1 · 1 citation
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Sound (cs.SD)
- On Addressing Practical Challenges for RNN-Transducer
2021/04/27 by Zhao, Rui, Xue, Jian, Li, Jinyu +3 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Continuous Speech Separation with Recurrent Selective Attention Network
2021/10/28 by Yixuan Zhang, Zhang, Yixuan, Zhuo Chen +10 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Speech Data
2022/03/31 by Junyi Ao, Ao, Junyi, Ziqiang Zhang +17 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Endpoint Detection for Streaming End-to-End Multi-talker ASR
2022/01/24 by Lu, Liang, Li, Jinyu, Gong, Yifan · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Total-Duration-Aware Duration Modeling for Text-to-Speech Systems
2024/06/06 by Şefik Emre Eskimez, Eskimez, Sefik Emre, Xiaofei Wang +21 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering
- SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training
2022/10/07 by Ziqiang Zhang, Long Zhou, Zhang, Ziqiang +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Self-supervised learning with bi-label masked speech prediction for streaming multi-talker speech recognition
2022/11/10 by Zili Huang, Huang, Zili, Zhuo Chen +14 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- CTCBERT: Advancing Hidden-unit BERT with CTC Objectives
2022/10/16 by Ruchao Fan, Yiming Wang, Fan, Ruchao +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Streaming, fast and accurate on-device Inverse Text Normalization for Automatic Speech Recognition
2022/11/07 by Yashesh Gaur, Nick Kibre, Gaur, Yashesh +13 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems
- Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
2025/04/14 by Yifan Yang, Yang, Yifan, Shujie Liu +22 · 7 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
2025/06/01 by Yao Qian, Zhang, Leying, Xiaofei Wang +17 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Building High-accuracy Multilingual ASR with Gated Language Experts and Curriculum Training
2023/03/01 by Eric Sun, Sun, Eric, Jinyu Li +19 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- t-SOT FNT: Streaming Multi-talker ASR with Text-only Domain Adaptation Capability
2023/09/15 by Wu, Jian, Kanda, Naoyuki, Yoshioka, Takuya +3 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- High-Accuracy and Low-Latency Speech Recognition with Two-Head Contextual Layer Trajectory LSTM Model
2020/03/17 by Jinyu Li, Li, Jinyu, Rui Zhao +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
2024/06/09 by Xiaofei Wang, Wang, Xiaofei, Şefik Emre Eskimez +19 · 1 citation
Engineering · Computer Science · #Advanced Adaptive Filtering Techniques #Speech and Audio Processing #Blind Source Separation Techniques
- Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
2024/12/20 by Y. F. Yang, Shujie Liu, Yang, Yifan +22 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering
- StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
2025/06/14 by Hui Wang, Wang, Hui, Shujie Liu +16 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Addressing speaker gender bias in large scale speech translation systems
2025/01/10 by Shubham Bansal, Vikas Joshi, Bansal, Shubham +7 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems