2024/05/24 by J. K. Becker, Jonas Becker, Becker, Jonas +7 · 9 citations
Computer Science · #A.1 #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2405.15604
Published in the Journal of Artificial Intelligence Research (JAIR)
openalex publication_date 2024/05/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28 · arxiv created 2026/08/06 · arxiv updated 2026/08/07
Text generation has become more accessible than ever, and the growing interest in these systems, especially those using large language models, has spurred a surge in related publications. We provide a systematic literature review comprising 257 papers, covering the period from January 2017 to December 2025. This review categorizes text generation contributions into five main tasks: open-ended text generation, summarization, translation, paraphrasing, and question answering. For each task in our taxonomy, we review relevant characteristics and key subtasks. We assess current approaches for evaluating text generation systems, covering model-free, model-based, and human evaluation. Our investigation shows several task-specific challenges (e.g., missing datasets for multi-document summarization, lack of coherence in story generation, and difficulties in complex reasoning for question answering). We further discuss nine challenges common to all tasks and sub-tasks in recent text generation papers: bias, reasoning, hallucinations, misuse, privacy, interpretability, transparency, datasets, and computing. This systematic literature review targets two main audiences: early-career researchers in natural language processing seeking an overview of the field and promising research directions, and senior researchers who need a recent overview of the main tasks, evaluation, challenges, and mitigation strategies.