2017/12/18 by Alexander Erdmann, Erdmann, Alexander, Nizar Habash +5
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.1712.06273
We present the second ever evaluated Arabic dialect-to-dialect machine\ntranslation effort, and the first to leverage external resources beyond a small\nparallel corpus. The subject has not previously received serious attention due\nto lack of naturally occurring parallel data; yet its importance is evidenced\nby dialectal Arabic's wide usage and breadth of inter-dialect variation,\ncomparable to that of Romance languages. Our results suggest that modeling\nmorphology and syntax significantly improves dialect-to-dialect translation,\nthough optimizing such data-sparse models requires consideration of the\nlinguistic differences between dialects and the nature of available data and\nresources. On a single-reference blind test set where untranslated input scores\n6.5 BLEU and a model trained only on parallel data reaches 14.6, pivot\ntechniques and morphosyntactic modeling significantly improve performance to\n17.5.\n