vix.ing · top · new · best · stats · spec

Curious Case of Language Generation Evaluation Metrics: A Cautionary\n Tale

2020/10/26 by Ozan Çağlayan, Pranava Madhyastha, Caglayan, Ozan +3 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2010.13588

openalex publication_date 2020/10/26 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

Automatic evaluation of language generation systems is a well-studied problem\nin Natural Language Processing. While novel metrics are proposed every year, a\nfew popular metrics remain as the de facto metrics to evaluate tasks such as\nimage captioning and machine translation, despite their known limitations. This\nis partly due to ease of use, and partly because researchers expect to see them\nand know how to interpret them. In this paper, we urge the community for more\ncareful consideration of how they automatically evaluate their models by\ndemonstrating important failure cases on multiple datasets, language pairs and\ntasks. Our experiments show that metrics (i) usually prefer system outputs to\nhuman-authored texts, (ii) can be insensitive to correct translations of rare\nwords, (iii) can yield surprisingly high scores when given a single sentence as\nsystem output for the entire test set.\n

Cited by

Related