vix.ing · top · new · best · stats · spec

Can Your Context-Aware MT System Pass the DiP Benchmark Tests? :\n Evaluation Benchmarks for Discourse Phenomena in Machine Translation

2020/04/30 by Prathyusha Jwalapuram, Barbara Rychalska, Jwalapuram, Prathyusha +5 · 4 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Text Readability and Simplification

paper · pdf · doi:10.48550/arxiv.2004.14607

Abstract

Despite increasing instances of machine translation (MT) systems including\ncontextual information, the evidence for translation quality improvement is\nsparse, especially for discourse phenomena. Popular metrics like BLEU are not\nexpressive or sensitive enough to capture quality improvements or drops that\nare minor in size but significant in perception. We introduce the first of\ntheir kind MT benchmark datasets that aim to track and hail improvements across\nfour main discourse phenomena: anaphora, lexical consistency, coherence and\nreadability, and discourse connective translation. We also introduce evaluation\nmethods for these tasks, and evaluate several baseline MT systems on the\ncurated datasets. Surprisingly, we find that existing context-aware models do\nnot improve discourse-related translations consistently across languages and\nphenomena.\n

Citations

Cited by

Related