2020/09/20 by Tahmid Hasan, Hasan, Tahmid, Abhik Bhattacharjee +11 · 6 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2009.09359
openalex publication_date 2020/09/20 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
Despite being the seventh most widely spoken language in the world, Bengali\nhas received much less attention in machine translation literature due to being\nlow in resources. Most publicly available parallel corpora for Bengali are not\nlarge enough; and have rather poor quality, mostly because of incorrect\nsentence alignments resulting from erroneous sentence segmentation, and also\nbecause of a high volume of noise present in them. In this work, we build a\ncustomized sentence segmenter for Bengali and propose two novel methods for\nparallel corpus creation on low-resource setups: aligner ensembling and batch\nfiltering. With the segmenter and the two methods combined, we compile a\nhigh-quality Bengali-English parallel corpus comprising of 2.75 million\nsentence pairs, more than 2 million of which were not available before.\nTraining on neural models, we achieve an improvement of more than 9 BLEU score\nover previous approaches to Bengali-English machine translation. We also\nevaluate on a new test set of 1000 pairs made with extensive quality control.\nWe release the segmenter, parallel corpus, and the evaluation set, thus\nelevating Bengali from its low-resource status. To the best of our knowledge,\nthis is the first ever large scale study on Bengali-English machine\ntranslation. We believe our study will pave the way for future research on\nBengali-English machine translation as well as other low-resource languages.\nOur data and code are available at https://github.com/csebuetnlp/banglanmt.\n