2020/06/26 by Aslı Çelikyılmaz, Asli Celikyilmaz, Elizabeth Clark +4 · 1 voice · 200 citations
Computer Science · Engineering · #Artificial intelligence #Automatic summarization #Computer science #Data science #Engineering #Focus (optics) #Information retrieval #Machine learning #Natural Language Processing Techniques #Natural language #Natural language generation #Natural language processing #Speech and dialogue systems #Systems engineering #Task (project management) #Text generation #Topic Modeling #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2006.14799
published in arXiv (Cornell University) (Cornell University) · 47 pages (revised version)
openalex publication_date 2020/06/26 · arxiv created 2021/05/18 · arxiv updated 2021/05/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/03
The paper surveys evaluation methods of natural language generation (NLG) systems that have been developed in the last few years. We group NLG evaluation methods into three categories: (1) human-centric evaluation metrics, (2) automatic metrics that require no training, and (3) machine-learned metrics. For each category, we discuss the progress that has been made and the challenges still being faced, with a focus on the evaluation of recently proposed NLG tasks and neural NLG models. We then present two examples for task-specific NLG evaluations for automatic text summarization and long text generation, and conclude the paper by proposing future research directions.