vix.ing · top · new · best · stats · spec

Between Flexibility and Consistency: Joint Generation of Captions and\n Subtitles

2021/07/13 by Alina Karakanta, Marco Gaido, Karakanta, Alina +5
Arts and Humanities · Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Subtitles and Audiovisual Media

paper · pdf · doi:10.48550/arxiv.2107.06246

openalex publication_date 2021/07/13 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

Speech translation (ST) has lately received growing interest for the\ngeneration of subtitles without the need for an intermediate source language\ntranscription and timing (i.e. captions). However, the joint generation of\nsource captions and target subtitles does not only bring potential output\nquality advantages when the two decoding processes inform each other, but it is\nalso often required in multilingual scenarios. In this work, we focus on ST\nmodels which generate consistent captions-subtitles in terms of structure and\nlexical content. We further introduce new metrics for evaluating subtitling\nconsistency. Our findings show that joint decoding leads to increased\nperformance and consistency between the generated captions and subtitles while\nstill allowing for sufficient flexibility to produce subtitles conforming to\nlanguage-specific needs and norms.\n

Related