Kumar, Rithesh
- High-Fidelity Audio Compression with Improved RVQGAN
2023/06/11 by Rithesh Kumar, Prem Seetharaman, Kumar, Rithesh +7 · 212 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
2019/10/08 by Kundan Kumar, Rithesh Kumar, Kumar, Kundan +15 · 132 citations
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
2016/12/22 by Mehri, Soroush, Kumar, Kundan, Gulrajani, Ishaan +5 · 10 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Sound (cs.SD)
- Chunked Autoregressive GAN for Conditional Waveform Synthesis
2021/10/19 by Max Morrison, Rithesh Kumar, Morrison, Max +9 · 11 citations
Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech and Audio Processing
- VampNet: Music Generation via Masked Acoustic Token Modeling
2023/07/10 by Hugo Flores García, Prem Seetharaman, Garcia, Hugo Flores +5 · 12 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Maximum Entropy Generators for Energy-Based Models
2019/01/24 by Rithesh Kumar, Kumar, Rithesh, Anirudh Goyal +6 · 8 citations
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- ObamaNet: Photo-realistic lip-sync from text
2017/12/06 by Rithesh Kumar, Jose Sotelo, Kumar, Rithesh +7 · 3 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Music and Audio Processing #Speech and Audio Processing
- NU-GAN: High resolution neural upsampling with GAN
2020/10/22 by Kumar, Rithesh, Kumar, Kundan, Anand, Vicki +2 · 3 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis
2024/10/14 by Y. Li, Li, Yingahao Aaron, Rithesh Kumar +3 · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
2025/04/13 by Guimarães, Heitor R., Su, Jiaqi, Kumar, Rithesh +2 · 5 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering