2014/10/04 by Valery Solovyev, Valery D. Solovyev, Solovyev, Valery D. +3
Arts and Humanities · Computer Science · Mathematics · Social Sciences · #62P25 #91F20 #Applications (stat.AP) #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #J.5 #Language and cultural evolution #Lexicography and Language Studies #Natural Language Processing Techniques #acm:62P25 #acm:91F20 #cs.CL #msc:62P25 #msc:91F20 #stat.AP
paper · pdf · doi:10.48550/arxiv.1410.1080
5 pages, 3 figures
arxiv created 2014/10/04 · openalex publication_date 2014/10/04 · arxiv updated 2014/10/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The article describes the original method of creating a dictionary of abbreviations based on the Google Books Ngram Corpus. The dictionary of abbreviations is designed for Russian, yet as its methodology is universal it can be applied to any language. The dictionary can be used to define the function of the period during text segmentation in various applied systems of text processing. The article describes difficulties encountered in the process of its construction as well as the ways to overcome them. A model of evaluating a probability of first and second type errors (extraction accuracy and fullness) is constructed. Certain statistical data for the use of abbreviations are provided.