2008/07/02 by Dmytro Lande, D. V. Lande, Lande, D. V. +2
Arts and Humanities · Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Lexicography and Language Studies #Literature, Language, and Rhetoric Studies #cs.CL #linguistics and terminology studies
paper · pdf · doi:10.48550/arxiv.0807.0311
3 pages
arxiv created 2008/07/02 · openalex publication_date 2008/07/02 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The algorithm of the creation texts parallel corpora was presented. The algorithm is based on the use of "key words" in text documents, and on the means of their automated translation. Key words were singled out by means of using Russian and Ukrainian morphological dictionaries, as well as dictionaries of the translation of nouns for the Russian and Ukrainianlanguages. Besides, to calculate the weights of the terms in the documents, empiric-statistic rules were used. The algorithm under consideration was realized in the form of a program complex, integrated into the content-monitoring InfoStream system. As a result, a parallel bilingual corpora of web-publications containing about 30 thousand documents, was created