2020/04/04 by Mihir Kale, Kale, Mihir, Scott Roy +1 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2004.02077
openalex publication_date 2020/04/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
While there is a large body of research studying deep learning methods for\ntext generation from structured data, almost all of it focuses purely on\nEnglish. In this paper, we study the effectiveness of machine translation based\npre-training for data-to-text generation in non-English languages. Since the\nstructured data is generally expressed in English, text generation into other\nlanguages involves elements of translation, transliteration and copying -\nelements already encoded in neural machine translation systems. Moreover, since\ndata-to-text corpora are typically small, this task can benefit greatly from\npre-training. Based on our experiments on Czech, a morphologically complex\nlanguage, we find that pre-training lets us train end-to-end models with\nsignificantly improved performance, as judged by automatic metrics and human\nevaluation. We also show that this approach enjoys several desirable\nproperties, including improved performance in low data scenarios and robustness\nto unseen slot values.\n