2020/04/28 by Rajesh Kumar Mundotiya, Manish Kumar Singh, Mundotiya, Rajesh Kumar +8 · 2 citations
Computer Science · Psychology · #Authorship Attribution and Profiling #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Second Language Acquisition and Learning #Text Readability and Simplification #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2004.13945
openalex publication_date 2020/04/28 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
Corpus preparation for low-resource languages and for development of human\nlanguage technology to analyze or computationally process them is a laborious\ntask, primarily due to the unavailability of expert linguists who are native\nspeakers of these languages and also due to the time and resources required.\nBhojpuri, Magahi, and Maithili, languages of the Purvanchal region of India (in\nthe north-eastern parts), are low-resource languages belonging to the\nIndo-Aryan (or Indic) family. They are closely related to Hindi, which is a\nrelatively high-resource language, which is why we compare with Hindi. We\ncollected corpora for these three languages from various sources and cleaned\nthem to the extent possible, without changing the data in them. The text\nbelongs to different domains and genres. We calculated some basic statistical\nmeasures for these corpora at character, word, syllable, and morpheme levels.\nThese corpora were also annotated with parts-of-speech (POS) and chunk tags.\nThe basic statistical measures were both absolute and relative and were\nexptected to indicate of linguistic properties such as morphological, lexical,\nphonological, and syntactic complexities (or richness). The results were\ncompared with a standard Hindi corpus. For most of the measures, we tried to\nthe corpus size the same across the languages to avoid the effect of corpus\nsize, but in some cases it turned out that using the full corpus was better,\neven if sizes were very different. Although the results are not very clear, we\ntry to draw some conclusions about the languages and the corpora. For POS\ntagging and chunking, the BIS tagset was used to manually annotate the data.\nThe POS tagged data sizes are 16067, 14669 and 12310 sentences, respectively,\nfor Bhojpuri, Magahi and Maithili. The sizes for chunking are 9695 and 1954\nsentences for Bhojpuri and Maithili, respectively.\n