2017/03/31 by Irshad Ahmad Bhat, Riyaz Ahmad Bhat, Bhat, Irshad Ahmad +5
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1703.10772
openalex publication_date 2017/03/31 · openalex created_date 2022/10/03 · openalex updated_date 2026/07/28
In this paper, we propose efficient and less resource-intensive strategies\nfor parsing of code-mixed data. These strategies are not constrained by\nin-domain annotations, rather they leverage pre-existing monolingual annotated\nresources for training. We show that these methods can produce significantly\nbetter results as compared to an informed baseline. Besides, we also present a\ndata set of 450 Hindi and English code-mixed tweets of Hindi multilingual\nspeakers for evaluation. The data set is manually annotated with Universal\nDependencies.\n