2017/07/18 by Elena Mikhalkova, Mikhalkova, Elena, Nadezhda Ganzherli +3
Computer Science · Physics and Astronomy · #Advanced Text Analysis Techniques #Complex Network Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Social and Information Networks (cs.SI) #Topic Modeling #cs.CL #cs.IR #cs.SI
paper · pdf · doi:10.48550/arxiv.1707.05481
11 pages, submitted for reviewing
openalex publication_date 2017/07/18 · arxiv created 2017/10/17 · arxiv updated 2017/10/18 · openalex created_date 2017/12/22 · openalex updated_date 2026/07/28
Being a matter of cognition, user interests should be apt to classification independent of the language of users, social network and content of interest itself. To prove it, we analyze a collection of English and Russian Twitter and Vkontakte community pages by interests of their followers. First, we create a model of Major Interests (MaIs) with the help of expert analysis and then classify a set of pages using machine learning algorithms (SVM, Neural Network, Naive Bayes, and some other). We take three interest domains that are typical of both English and Russian-speaking communities: football, rock music, vegetarianism. The results of classification show a greater correlation between Russian-Vkontakte and Russian-Twitter pages while English-Twitterpages appear to provide the highest score.