2018/09/06 by Saurabh Garg, Garg, Saurabh, Tanmay Parekh +3
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1809.01962
openalex publication_date 2018/09/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This work focuses on building language models (LMs) for code-switched text.\nWe propose two techniques that significantly improve these LMs: 1) A novel\nrecurrent neural network unit with dual components that focus on each language\nin the code-switched text separately 2) Pretraining the LM using synthetic text\nfrom a generative model estimated using the training data. We demonstrate the\neffectiveness of our proposed techniques by reporting perplexities on a\nMandarin-English task and derive significant reductions in perplexity.\n