vix.ing · top · new · best · stats

Sentence-Level Fluency Evaluation: References Help, But Can Be Spared!

2018/09/24 by Katharina Kann, Kann, Katharina, Sascha Rothe +3 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling #cs.CL

paper · pdf · doi:10.48550/arxiv.1809.08731

Accepted to CoNLL 2018

arxiv created 2018/09/24 · openalex publication_date 2018/09/24 · arxiv updated 2018/09/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Motivated by recent findings on the probabilistic modeling of acceptability judgments, we propose syntactic log-odds ratio (SLOR), a normalized language model score, as a metric for referenceless fluency evaluation of natural language generation output at the sentence level. We further introduce WPSLOR, a novel WordPiece-based version, which harnesses a more compact language model. Even though word-overlap metrics like ROUGE are computed with the help of hand-written references, our referenceless methods obtain a significantly higher correlation with human fluency scores on a benchmark dataset of compressed sentences. Finally, we present ROUGE-LM, a reference-based metric which is a natural extension of WPSLOR to the case of available references. We show that ROUGE-LM yields a significantly higher correlation with human judgments than all baseline metrics, including WPSLOR on its own.

Citations

Cited by

Related