vix.ing · top · new · best · stats · spec

Yan, Zhijie

  1. Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
    2023/11/14 by Yunfei Chu, Chu, Yunfei, Jin Xu +13 · 178 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  2. CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
    2024/07/07 by Zhihao Du, Du, Zhihao, Qian Chen +21 · 127 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Video Analysis and Summarization #electronic engineering #information engineering
  3. CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
    2024/12/13 by Zhihao Du, Yuxuan Wang, Du, Zhihao +35 · 140 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
  4. Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
    2022/06/16 by Zhifu Gao, Shiliang Zhang, Gao, Zhifu +5 · 47 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. M2MeT: The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
    2021/10/14 by Fan Yu, Shiliang Zhang, Yu, Fan +21 · 22 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
  6. FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
    2024/07/04 by An, Keyu, Chen, Qian, Deng, Chong +30 · 32 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  7. MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
    2025/01/10 by Chen, Qian, Chen, Yafeng, Chen, Yanni +33 · 30 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #electronic engineering #information engineering
  8. Dynamic Open-Vocabulary 3D Scene Graphs for Long-term Language-Guided Mobile Manipulation
    2024/10/15 by Zhijie Yan, Yan, Zhijie, Shufei Li +13 · 16 citations
    Computer Science · Engineering · #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Robot Manipulation and Learning #Robotics (cs.RO)
  9. LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
    2023/10/07 by Zhihao Du, Jiaming Wang, Du, Zhihao +27 · 10 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  10. TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
    2024/03/28 by Bu Jin, Jin, Bu, Yupeng Zheng +27 · 11 citations
    Computer Science · #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Video Analysis and Summarization
  11. Deep-FSMN for Large Vocabulary Continuous Speech Recognition
    2018/03/04 by Shiliang Zhang, Zhang, Shiliang, Ming Lei +5 · 6 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  12. Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
    2022/02/08 by Fan Yu, Yu, Fan, Shiliang Zhang +29 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  13. Large Language Models Powered Context-aware Motion Prediction in Autonomous Driving
    2024/03/17 by Xiaoji Zheng, Zheng, Xiaoji, Lixiu Wu +13 · 5 citations
    Computer Science · #68T45 #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Robotics (cs.RO)
  14. OmniAudio: Generating Spatial Audio from 360-Degree Video
    2025/04/21 by Liu, Huadai, Luo, Tianyi, Luo, Kaicheng +11 · 9 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  15. The second multi-channel multi-party meeting transcription challenge (M2MeT) 2.0): A benchmark for speaker-attributed ASR
    2023/09/24 by Yuhao Liang, Mohan Shi, Liang, Yuhao +24 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  16. Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis
    2022/11/18 by Du, Zhihao, Zhang, Shiliang, Zheng, Siqi +1 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  17. SyncSpeech: Low-Latency and Efficient Dual-Stream Text-to-Speech based on Temporal Masked Transformer
    2025/02/16 by Sheng, Zhengyan, Du, Zhihao, Zhang, Shiliang +3 · 5 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Sound (cs.SD)
  18. ProsoSpeech: Enhancing Prosody With Quantized Vector Pre-training in Text-to-Speech
    2022/02/16 by Ren, Yi, Lei, Ming, Huang, Zhiying +4 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  19. Achieving Timestamp Prediction While Recognizing with Non-Autoregressive End-to-End ASR Model
    2023/01/29 by Xian Shi, Yanni Chen, Shi, Xian +5 · 1 citation
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #EEG and Brain-Computer Interfaces #FOS: Computer and information sciences #FOS: Electrical engineering #Gaze Tracking and Assistive Technology #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  20. Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
    2024/09/26 by Keyu An, Shiliang Zhang, An, Keyu +3 · 1 citation
    Engineering · Computer Science · #Fault Detection and Control Systems #Sensor Technology and Measurement Systems #Neural Networks and Applications