2021/06/30 by Elizabeth Clark, Elizabeth A. Clark, Clark, Elizabeth +10 · 57 citations
Computer Science · Mathematics · Psychology · #Artificial intelligence #Computer science #Explainable Artificial Intelligence (XAI) #Fluency #Gold standard (test) #Mathematics #Mathematics education #Natural Language Processing Techniques #Natural language #Natural language generation #Natural language processing #Psychology #Statistics #Text generation #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2107.00061
published in arXiv (Cornell University) (Cornell University) · references added, corrected typo
openalex publication_date 2021/06/30 · arxiv created 2021/07/07 · arxiv updated 2021/07/08 · openalex created_date 2022/07/25 · openalex updated_date 2026/08/06
Human evaluations are typically considered the gold standard in natural\nlanguage generation, but as models' fluency improves, how well can evaluators\ndetect and judge machine-generated text? We run a study assessing non-experts'\nability to distinguish between human- and machine-authored text (GPT2 and GPT3)\nin three domains (stories, news articles, and recipes). We find that, without\ntraining, evaluators distinguished between GPT3- and human-authored text at\nrandom chance level. We explore three approaches for quickly training\nevaluators to better identify GPT3-authored text (detailed instructions,\nannotated examples, and paired examples) and find that while evaluators'\naccuracy improved up to 55%, it did not significantly improve across the three\ndomains. Given the inconsistent results across text domains and the often\ncontradictory reasons evaluators gave for their judgments, we examine the role\nuntrained human evaluations play in NLG evaluation and provide recommendations\nto NLG researchers for improving human evaluations of text generated from\nstate-of-the-art models.\n