Yan, Zhijie
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
2023/11/14 by Yunfei Chu, Chu, Yunfei, Jin Xu +13 · 178 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
2024/07/07 by Zhihao Du, Du, Zhihao, Qian Chen +21 · 127 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Video Analysis and Summarization #electronic engineering #information engineering
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
2024/12/13 by Zhihao Du, Yuxuan Wang, Du, Zhihao +35 · 140 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
- Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
2022/06/16 by Zhifu Gao, Shiliang Zhang, Gao, Zhifu +5 · 47 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- M2MeT: The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
2021/10/14 by Fan Yu, Shiliang Zhang, Yu, Fan +21 · 22 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
- FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
2024/07/04 by An, Keyu, Chen, Qian, Deng, Chong +30 · 32 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
- MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
2025/01/10 by Chen, Qian, Chen, Yafeng, Chen, Yanni +33 · 30 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #electronic engineering #information engineering
- Dynamic Open-Vocabulary 3D Scene Graphs for Long-term Language-Guided Mobile Manipulation
2024/10/15 by Zhijie Yan, Yan, Zhijie, Shufei Li +13 · 16 citations
Computer Science · Engineering · #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Robot Manipulation and Learning #Robotics (cs.RO)
- LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
2023/10/07 by Zhihao Du, Jiaming Wang, Du, Zhihao +27 · 10 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
2024/03/28 by Bu Jin, Jin, Bu, Yupeng Zheng +27 · 11 citations
Computer Science · #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Video Analysis and Summarization
- Deep-FSMN for Large Vocabulary Continuous Speech Recognition
2018/03/04 by Shiliang Zhang, Zhang, Shiliang, Ming Lei +5 · 6 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
2022/02/08 by Fan Yu, Yu, Fan, Shiliang Zhang +29 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- Large Language Models Powered Context-aware Motion Prediction in Autonomous Driving
2024/03/17 by Xiaoji Zheng, Zheng, Xiaoji, Lixiu Wu +13 · 5 citations
Computer Science · #68T45 #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Robotics (cs.RO)
- OmniAudio: Generating Spatial Audio from 360-Degree Video
2025/04/21 by Liu, Huadai, Luo, Tianyi, Luo, Kaicheng +11 · 9 citations
#Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- The second multi-channel multi-party meeting transcription challenge (M2MeT) 2.0): A benchmark for speaker-attributed ASR
2023/09/24 by Yuhao Liang, Mohan Shi, Liang, Yuhao +24 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis
2022/11/18 by Du, Zhihao, Zhang, Shiliang, Zheng, Siqi +1 · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
- SyncSpeech: Low-Latency and Efficient Dual-Stream Text-to-Speech based on Temporal Masked Transformer
2025/02/16 by Sheng, Zhengyan, Du, Zhihao, Zhang, Shiliang +3 · 5 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Sound (cs.SD)
- ProsoSpeech: Enhancing Prosody With Quantized Vector Pre-training in Text-to-Speech
2022/02/16 by Ren, Yi, Lei, Ming, Huang, Zhiying +4 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Achieving Timestamp Prediction While Recognizing with Non-Autoregressive End-to-End ASR Model
2023/01/29 by Xian Shi, Yanni Chen, Shi, Xian +5 · 1 citation
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #EEG and Brain-Computer Interfaces #FOS: Computer and information sciences #FOS: Electrical engineering #Gaze Tracking and Assistive Technology #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
- Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
2024/09/26 by Keyu An, Shiliang Zhang, An, Keyu +3 · 1 citation
Engineering · Computer Science · #Fault Detection and Control Systems #Sensor Technology and Measurement Systems #Neural Networks and Applications