vix.ing · top · new · best · stats · spec

Jaehyeon Kim

  1. HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
    2020/10/12 by Jungil Kong, Kong, Jungil, Jaehyeon Kim +3 · 182 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  2. Conditional Variational Autoencoder with Adversarial Learning for\n End-to-End Text-to-Speech
    2021/06/10 by Jaehyeon Kim, Kim, Jaehyeon, Jungil Kong +3 · 78 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search
    2020/05/22 by Jaehyeon Kim, Sungwon Kim, Kim, Jaehyeon +5 · 30 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  4. CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
    2024/04/03 by Jaehyeon Kim, Keon Lee, Kim, Jaehyeon +5 · 10 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  5. DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
    2024/06/17 by Keon Lee, Lee, Keon, Dong-Won Kim +5 · 6 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  6. Efficient Generative Modeling with Residual Vector Quantization-Based Tokens
    2024/12/13 by Jaehyeon Kim, T. Moon, Kim, Jaehyeon +5 · 4 citations
    Computer Science · #Algorithms and Data Compression #Distributed and Parallel Computing Systems #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications
  7. How to Move Your Dragon: Text-to-Motion Synthesis for Large-Vocabulary Objects
    2025/03/06 by Wonkwang Lee, Lee, Wonkwang, Jongwon Jeong +11 · 2 citations
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation
  8. PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
    2026/01/14 by Rajarshi Roy, Jonathan Raiman, Sang-gil Lee +5 · 1 voice · 1 citation
    Computer Science · #cs.CL
  9. Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
    2026/07/17 by Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar +19
    #eess.AS #cs.CV