2020/06/11 by Nitika Mathur, Timothy Baldwin, Mathur, Nitika +3 · 12 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Software Engineering Research #Software System Performance and Reliability #Software Testing and Debugging Techniques
paper · pdf · doi:10.48550/arxiv.2006.06264
openalex publication_date 2020/06/11 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
Automatic metrics are fundamental for the development and evaluation of\nmachine translation systems. Judging whether, and to what extent, automatic\nmetrics concur with the gold standard of human evaluation is not a\nstraightforward problem. We show that current methods for judging metrics are\nhighly sensitive to the translations used for assessment, particularly the\npresence of outliers, which often leads to falsely confident conclusions about\na metric's efficacy. Finally, we turn to pairwise system ranking, developing a\nmethod for thresholding performance improvement under an automatic metric\nagainst human judgements, which allows quantification of type I versus type II\nerrors incurred, i.e., insignificant human differences in system quality that\nare accepted, and significant human differences that are rejected. Together,\nthese findings suggest improvements to the protocols for metric evaluation and\nsystem performance evaluation in machine translation.\n