2021/11/09 by Antoine Caillon, Caillon, Antoine, Philippe Esling +1 · 22 citations
Computer Science · Engineering · #Artificial intelligence #Audio and Speech Processing (eess.AS) #Audio signal #Autoencoder #Computer science #Deep learning #Digital audio #FOS: Computer and information sciences #FOS: Electrical engineering #Generative grammar #Generative model #High fidelity #Machine Learning (cs.LG) #Machine learning #Music Technology and Sound Studies #Music and Audio Processing #Pattern recognition (psychology) #Representation (politics) #Sound (cs.SD) #Sound quality #Speech Recognition and Synthesis #Speech and Audio Processing #Speech coding #Speech recognition #Waveform #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2111.05011
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2021/11/09 · arxiv created 2021/12/15 · arxiv updated 2021/12/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/08
Deep generative models applied to audio have improved by a large margin the state-of-the-art in many speech and music related tasks. However, as raw waveform modelling remains an inherently difficult task, audio generative models are either computationally intensive, rely on low sampling rates, are complicated to control or restrict the nature of possible signals. Among those models, Variational AutoEncoders (VAE) give control over the generation by exposing latent variables, although they usually suffer from low synthesis quality. In this paper, we introduce a Realtime Audio Variational autoEncoder (RAVE) allowing both fast and high-quality audio waveform synthesis. We introduce a novel two-stage training procedure, namely representation learning and adversarial fine-tuning. We show that using a post-training analysis of the latent space allows a direct control between the reconstruction fidelity and the representation compactness. By leveraging a multi-band decomposition of the raw waveform, we show that our model is the first able to generate 48kHz audio signals, while simultaneously running 20 times faster than real-time on a standard laptop CPU. We evaluate synthesis quality using both quantitative and qualitative subjective experiments and show the superiority of our approach compared to existing models. Finally, we present applications of our model for timbre transfer and signal compression. All of our source code and audio examples are publicly available.