vix.ing · top · new · best · stats · spec

Zuo, Jialong

  1. WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
    2024/08/29 by Shengpeng Ji, Ziyue Karen Jiang, Ji, Shengpeng +30 · 38 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  2. WavChat: A Survey of Spoken Dialogue Models
    2024/11/15 by Shengpeng Ji, Yifu Chen, Ji, Shengpeng +35 · 21 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multi-Agent Systems and Negotiation #Multimedia (cs.MM) #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  3. Holmes-VAD: Towards Unbiased and Explainable Video Anomaly Detection via Multi-modal LLM
    2024/06/18 by Zhang, Huaxin, Xu, Xiaohao, Wang, Xiang +6 · 14 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  4. PLIP: Language-Image Pre-training for Person Representation Learning
    2023/05/15 by Jialong Zuo, Zuo, Jialong, Hong, Jiahao +9 · 8 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gait Recognition and Analysis #Human Pose and Action Recognition #Video Surveillance and Tracking Methods
  5. Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
    2024/12/09 by Hua‐Xin Zhang, Xiaohao Xu, Zhang, Huaxin +15 · 16 citations
    Computer Science · #Anomaly Detection Techniques and Applications #Digital Media Forensic Detection #Generative Adversarial Networks and Image Synthesis
  6. MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
    2025/02/26 by Ziyue Karen Jiang, Jiang, Ziyue, Yi Ren +24 · 19 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
  7. ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
    2024/06/03 by Shengpeng Ji, Qian Chen, Ji, Shengpeng +18 · 8 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  8. FluentSpeech: Stutter-Oriented Automatic Speech Editing with Context-Aware Diffusion Models
    2023/05/23 by Ziyue Karen Jiang, Jiang, Ziyue, Qian Yang +11 · 5 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
    2024/10/28 by Cheng, Xize, Zheng, Siqi, Wang, Zehan +8 · 6 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  10. Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
    2024/02/19 by Ji, Shengpeng, Fang, Minghui, Zuo, Jialong +5 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  11. UFineBench: Towards Text-based Person Retrieval with Ultra-fine Granularity
    2023/12/06 by Jialong Zuo, Zuo, Jialong, Hanyu Zhou +13 · 2 citations
    Computer Science · Social Sciences · #Video Surveillance and Tracking Methods #Human Pose and Action Recognition #Human Mobility and Location-Based Analysis
  12. Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
    2025/02/08 by Jialong Zuo, Shengpeng Ji, Zuo, Jialong +19 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
  13. CART: A Generative Cross-Modal Retrieval Framework with Coarse-To-Fine Semantic Modeling
    2024/06/25 by Fang, Minghui, Ji, Shengpeng, Zuo, Jialong +9 · 1 citation
    #FOS: Computer and information sciences #Information Retrieval (cs.IR)
  14. WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
    2025/05/14 by Shengpeng Ji, Ji, Shengpeng, Liang, Tianle +22 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multi-Agent Systems and Negotiation #Multimedia (cs.MM) #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  15. VideoLucy: Deep Memory Backtracking for Long Video Understanding
    2025/10/14 by Zuo, Jialong, Deng, Yongtai, Kong, Lingdong +7 · 3 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  16. MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
    2024/07/19 by Yang, Qian, Zuo, Jialong, Su, Zhe +6 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  17. MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
    2024/02/14 by Ji, Shengpeng, Jiang, Ziyue, Wang, Hanting +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering