2025/01/09 by Tanel Alumäe, Fedorchenko, Artem, Alumäe, Tanel · 1 citation
Arts and Humanities · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Subtitles and Audiovisual Media #Translation Studies and Practices #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2501.05234
openalex publication_date 2025/01/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper presents an approach for generating high-quality, same-language\nsubtitles for Estonian TV content. We fine-tune the Whisper model on\nhuman-generated Estonian subtitles and enhance it with iterative\npseudo-labeling and large language model (LLM) based post-editing. Our\nexperiments demonstrate notable subtitle quality improvement through\npseudo-labeling with an unlabeled dataset. We find that applying LLM-based\nediting at test time enhances subtitle accuracy, while its use during training\ndoes not yield further gains. This approach holds promise for creating subtitle\nquality close to human standard and could be extended to real-time\napplications.\n