vix.ing · top · new · best · stats · spec

Zhong, Zhi

  1. SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
    2024/06/25 by Marco Comunità, Zhi Zhong, Comunità, Marco +17 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  2. SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
    2024/05/28 by Koichi Saito, Saito, Koichi, D. S. Kim +11 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
  3. Extending Audio Masked Autoencoders Toward Audio Restoration
    2023/05/11 by Zhi Zhong, Hao Shi, Zhong, Zhi +13 · 2 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Ultrasonics and Acoustic Wave Propagation #electronic engineering #information engineering
  4. Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
    2023/05/18 by Shi, Hao, Shimada, Kazuki, Hirano, Masato +6 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  5. Music Foundation Model as Generic Booster for Music Downstream Tasks
    2024/11/02 by Liao, WeiHsiang, Takida, Yuhta, Ikemiya, Yukara +13 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  6. Variable Bitrate Residual Vector Quantization for Audio Coding
    2024/10/08 by Chae, Yunkee, Choi, Woosung, Takida, Yuhta +8 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  7. Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
    2024/05/23 by Shiqi Yang, Zhi Zhong, Yang, Shiqi +11 · 1 citation
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  8. OpenMU: Your Swiss Army Knife for Music Understanding
    2024/10/21 by Mengjie Zhao, Zhi Zhong, Zhao, Mengjie +13 · 1 citation
    Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  9. Cross-Modal Learning for Music-to-Music-Video Description Generation
    2025/03/14 by Zhuoyuan Mao, Mao, Zhuoyuan, Mengjie Zhao +11 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering
  10. SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
    2025/05/22 by Zhi Zhong, Zhong, Zhi, Akira Takahashi +9 · 4 citations
    Computer Science · #Music Technology and Sound Studies #Music and Audio Processing #Generative Adversarial Networks and Image Synthesis