2019/09/10 by Binny Mathew, Mathew, Binny, Suman Kalyan Maity +5
Computer Science · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #Expert finding and Q&A systems #FOS: Computer and information sciences #Information Retrieval and Search Behavior #Social and Information Networks (cs.SI) #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1909.04367
openalex publication_date 2019/09/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Quora is a popular Q&A site which provides users with the ability to tag\nquestions with multiple relevant topics which helps to attract quality answers.\nThese topics are not predefined but user-defined conventions and it is not so\nrare to have multiple such conventions present in the Quora ecosystem\ndescribing exactly the same concept. In almost all such cases, users (or Quora\nmoderators) manually merge the topic pair into one of the either topics, thus\nselecting one of the competing conventions. An important application for the\nsite therefore is to identify such competing conventions early enough that\nshould merge in future. In this paper, we propose a two-step approach that\nuniquely combines the anomaly detection and the supervised classification\nframeworks to predict whether two topics from among millions of topic pairs are\nindeed competing conventions, and should merge, achieving an F-score of 0.711.\nWe also develop a model to predict the direction of the topic merge, i.e., the\nwinning convention, achieving an F-score of 0.898. Our system is also able to\npredict ~ 25% of the correct case of merges within the first month of the merge\nand ~ 40% of the cases within a year. This is an encouraging result since Quora\nusers on average take 936 days to identify such a correct merge. Human judgment\nexperiments show that our system is able to predict almost all the correct\ncases that humans can predict plus 37.24% correct cases which the humans are\nnot able to identify at all.\n