vix.ing · top · new · best · stats · spec

Phrase Based Language Model for Statistical Machine Translation: Empirical Study

2015/01/21 by Geliang Chen, Chen, Geliang
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text and Document Classification Technologies #Topic Modeling #cs.CL

paper · pdf · doi:10.48550/arxiv.1501.05203

supplementary material of http://arxiv.org/abs/1501.04324. This version is identical to the Bachelor thesis of Geliang Chen archived on the 20th June 2013 in Peking University. Thesis advisor: Professor Jia Xu

openalex publication_date 2015/01/21 · arxiv created 2015/02/18 · arxiv updated 2015/02/19 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28

Abstract

Reordering is a challenge to machine translation (MT) systems. In MT, the widely used approach is to apply word based language model (LM) which considers the constituent units of a sentence as words. In speech recognition (SR), some phrase based LM have been proposed. However, those LMs are not necessarily suitable or optimal for reordering. We propose two phrase based LMs which considers the constituent units of a sentence as phrases. Experiments show that our phrase based LMs outperform the word based LM with the respect of perplexity and n-best list re-ranking.

Citations

Related