vix.ing · top · new · best · stats · spec

AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for\n Indic Languages

2020/04/30 by Anoop Kunchukuttan, Divyanshu Kakwani, Kunchukuttan, Anoop +12 · 5 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Text Readability and Simplification

paper · pdf · doi:10.48550/arxiv.2005.00085

Abstract

We present the IndicNLP corpus, a large-scale, general-domain corpus\ncontaining 2.7 billion words for 10 Indian languages from two language\nfamilies. We share pre-trained word embeddings trained on these corpora. We\ncreate news article category classification datasets for 9 languages to\nevaluate the embeddings. We show that the IndicNLP embeddings significantly\noutperform publicly available pre-trained embedding on multiple evaluation\ntasks. We hope that the availability of the corpus will accelerate Indic NLP\nresearch. The resources are available at\nhttps://github.com/ai4bharat-indicnlp/indicnlpcorpus.\n

Cited by

Related