vix.ing · top · new · best · stats · spec

Kong, Zhifeng

  1. DiffWave: A Versatile Diffusion Model for Audio Synthesis
    2020/09/21 by Kong, Zhifeng, Ping, Wei, Huang, Jiaji +2 · 116 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
  2. Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
    2024/02/02 by Zhifeng Kong, Arushi Goel, Kong, Zhifeng +8 · 41 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  3. Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
    2025/03/06 by Ghosh, Sreyan, Kong, Zhifeng, Kumar, Sonal +6 · 44 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  4. On Fast Sampling of Diffusion Probabilistic Models
    2021/05/31 by Zhifeng Kong, Kong, Zhifeng, Wei Ping +1 · 14 citations
    Computer Science · Medicine · Mathematics · #Generative Adversarial Networks and Image Synthesis #Advanced Neuroimaging Techniques and Applications #Markov Chains and Monte Carlo Methods
  5. A Conditional Point Diffusion-Refinement Paradigm for 3D Point Cloud Completion
    2021/12/07 by Lyu, Zhaoyang, Kong, Zhifeng, Xu, Xudong +2 · 11 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  6. TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
    2024/12/30 by Hung, Chia-Yu, Majumder, Navonil, Kong, Zhifeng +6 · 25 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  7. Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
    2025/07/10 by Goel, Arushi, Ghosh, Sreyan, Kim, Jaehyeon +8 · 59 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  8. ETTA: Elucidating the Design Space of Text-to-Audio Models
    2024/12/26 by Lee, Sang-gil, Kong, Zhifeng, Goel, Arushi +3 · 6 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  9. Speech Denoising in the Waveform Domain with Self-Attention
    2022/02/15 by Zhifeng Kong, Kong, Zhifeng, Ping, Wei +4 · 2 citations
    Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. Understanding Instance-based Interpretability of Variational Auto-Encoders
    2021/05/29 by Zhifeng Kong, Kamalika Chaudhuri, Kong, Zhifeng +1 · 2 citations
    Computer Science · Physics and Astronomy · #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Model Reduction and Neural Networks
  11. A2SB: Audio-to-Audio Schrodinger Bridges
    2025/01/20 by Kong, Zhifeng, Shih, Kevin J, Nie, Weili +6 · 4 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  12. Fastened CROWN: Tightened Neural Network Robustness Certificates
    2019/12/02 by Lyu, Zhaoyang, Ko, Ching-Yun, Kong, Zhifeng +3 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  13. Improving Text-To-Audio Models with Synthetic Captions
    2024/06/18 by Kong, Zhifeng, Lee, Sang-gil, Ghosal, Deepanway +5 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  14. A Geometry-Aware Algorithm to Learn Hierarchical Embeddings in Hyperbolic Space
    2024/07/23 by Zhangyu Wang, Lantian Xu, Wang, Zhangyu +9 · 2 citations
    Computer Science · Engineering · #Image Processing and 3D Reconstruction #3D Shape Modeling and Analysis #Human Motion and Animation
  15. Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic Data
    2024/10/02 by Ghosh, Sreyan, Kumar, Sonal, Kong, Zhifeng +3 · 2 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  16. The Expressive Power of a Class of Normalizing Flow Models
    2020/05/31 by Zhifeng Kong, Kamalika Chaudhuri, Kong, Zhifeng +1 · 1 citation
    Computer Science · #Algorithms and Data Compression #Artificial Intelligence in Games #Cellular Automata and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  17. Audio Dialogues: Dialogues dataset for audio and music understanding
    2024/04/11 by Arushi Goel, Goel, Arushi, Zhifeng Kong +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  18. Multi-Domain Audio Question Answering Toward Acoustic Content Reasoning in The DCASE 2025 Challenge
    2025/05/12 by Yang, Chao-Han Huck, Ghosh, Sreyan, Wang, Qing +14 · 2 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering