2020/10/23 by Oshin Agarwal, Heming Ge, Agarwal, Oshin +5
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2010.12688
openalex publication_date 2020/10/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Prior work on Data-To-Text Generation, the task of converting knowledge graph\n(KG) triples into natural text, focused on domain-specific benchmark datasets.\nIn this paper, however, we verbalize the entire English Wikidata KG, and\ndiscuss the unique challenges associated with a broad, open-domain, large-scale\nverbalization. We further show that verbalizing a comprehensive, encyclopedic\nKG like Wikidata can be used to integrate structured KGs and natural language\ncorpora. In contrast to the many architectures that have been developed to\nintegrate these two sources, our approach converts the KG into natural text,\nallowing it to be seamlessly integrated into existing language models. It\ncarries the further advantages of improved factual accuracy and reduced\ntoxicity in the resulting language model. We evaluate this approach by\naugmenting the retrieval corpus in a retrieval language model and showing\nsignificant improvements on the knowledge intensive tasks of open domain QA and\nthe LAMA knowledge probe.\n