vix.ing · top · new · best · stats · spec

Ziyue Karen Jiang

  1. AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
    2024/02/12 by Qian Yang, Yang, Qian, Jin Xu +19 · 62 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  2. WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
    2024/08/29 by Shengpeng Ji, Ji, Shengpeng, Ziyue Karen Jiang +30 · 41 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis
    2023/01/31 by Zhenhui Ye, Ye, Zhenhui, Ziyue Karen Jiang +9 · 20 citations
    Computer Science · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Speech and Audio Processing
  4. GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation
    2023/05/01 by Zhenhui Ye, Jinzheng He, Ye, Zhenhui +17 · 8 citations
    Computer Science · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Speech and Audio Processing
  5. MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
    2025/02/26 by Ziyue Karen Jiang, Jiang, Ziyue, Yi Ren +24 · 20 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
  6. Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias
    2023/06/06 by Ziyue Karen Jiang, Jiang, Ziyue, Yi Ren +21 · 7 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
    2023/07/14 by Ziyue Karen Jiang, Jinglin Liu, Jiang, Ziyue +21 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
    2024/06/03 by Shengpeng Ji, Ji, Shengpeng, Qian Chen +18 · 8 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  9. FluentSpeech: Stutter-Oriented Automatic Speech Editing with Context-Aware Diffusion Models
    2023/05/23 by Ziyue Karen Jiang, Jiang, Ziyue, Qian Yang +11 · 5 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. Make-A-Voice: Unified Voice Synthesis With Discrete Representation
    2023/05/30 by Rongjie Huang, Huang, Rongjie, Chunlei Zhang +17 · 5 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
  11. CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-training
    2023/05/18 by Zhenhui Ye, Rongjie Huang, Ye, Zhenhui +13 · 4 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  12. Ada-TTA: Towards Adaptive High-Quality Text-to-Talking Avatar Synthesis
    2023/06/06 by Zhenhui Ye, Ye, Zhenhui, Ziyue Karen Jiang +13 · 2 citations
    Computer Science · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Speech and Audio Processing
  13. Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
    2025/02/08 by Jialong Zuo, Shengpeng Ji, Zuo, Jialong +19 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
  14. Versatile Framework for Song Generation with Prompt-based Control
    2025/04/27 by Yu Zhang, Zhang, Yu, Wenxiang Guo +19 · 1 citation
    Computer Science · #Artificial Intelligence in Games #Music Technology and Sound Studies #Music and Audio Processing