2015/06/23 by Michel Galley, Galley, Michel, Chris Brockett +15 · 12 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.1506.06863
6 pages, to appear at ACL 2015
openalex publication_date 2015/06/23 · arxiv created 2015/06/24 · arxiv updated 2015/06/25 · openalex created_date 2022/10/03 · openalex updated_date 2026/07/28
We introduce Discriminative BLEU (deltaBLEU), a novel metric for intrinsic evaluation of generated text in tasks that admit a diverse range of possible outputs. Reference strings are scored for quality by human raters on a scale of [-1, +1] to weight multi-reference BLEU. In tasks involving generation of conversational responses, deltaBLEU correlates reasonably with human judgments and outperforms sentence-level and IBM BLEU in terms of both Spearman's rho and Kendall's tau.