2016/10/26 by Dimitra Gkatzia, Gkatzia, Dimitra
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Semantic Web and Ontologies #Topic Modeling #Web Data Mining and Analysis
paper · pdf · doi:10.48550/arxiv.1610.08375
openalex publication_date 2016/10/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Data-to-text systems are powerful in generating reports from data automatically and thus they simplify the presentation of complex data. Rather than presenting data using visualisation techniques, data-to-text systems use natural (human) language, which is the most common way for human-human communication. In addition, data-to-text systems can adapt their output content to users' preferences, background or interests and therefore they can be pleasant for users to interact with. Content selection is an important part of every data-to-text system, because it is the module that determines which from the available information should be conveyed to the user. This survey initially introduces the field of data-to-text generation, describes the general data-to-text system architecture and then it reviews the state-of-the-art content selection methods. Finally, it provides recommendations for choosing an approach and discusses opportunities for future research.