2024/09/02 by Jaeyeon Kim, Kim, Jaeyeon, Minjeon Jeon +7 · 1 citation
Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Subtitles and Audiovisual Media #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2409.01201
openalex publication_date 2024/09/02 · openalex created_date 2024/09/29 · openalex updated_date 2026/07/28
In this work, we aim to analyze and optimize the EnCLAP framework, a state-of-the-art model in automated audio captioning. We investigate the impact of modifying the acoustic encoder components, explore pretraining with different dataset scales, and study the effectiveness of a reranking scheme. Through extensive experimentation and quantitative analysis of generated captions, we develop EnCLAP++, an enhanced version that significantly surpasses the original.