2019/02/06 by Xiang Fu, Fu, Xiang, Shangdi Yu +3
Computer Science · Mathematics · Physics and Astronomy · #Complex Network Analysis Techniques #Expert finding and Q&A systems #FOS: Computer and information sciences #FOS: Physical sciences #Machine Learning (stat.ML) #Physics and Society (physics.soc-ph) #Social and Information Networks (cs.SI) #Topic Modeling #cs.SI #physics.soc-ph #stat.ML
paper · pdf · doi:10.48550/arxiv.1902.02372
arxiv created 2019/02/06 · openalex publication_date 2019/02/06 · arxiv updated 2019/02/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Large Question-and-Answer (Q&A) platforms support diverse knowledge curation on the Web. While researchers have studied user behavior on the platforms in a variety of contexts, there is relatively little insight into important by-products of user behavior that also encode knowledge. Here, we analyze and model the macroscopic structure of tags applied by users to annotate and catalog questions, using a collection of 168 Stack Exchange websites. We find striking similarity in tagging structure across these Stack Exchange communities, even though each community evolves independently (albeit under similar guidelines). Using our empirical findings, we develop a simple generative model that creates random bipartite graphs of tags and questions. Our model accounts for the tag frequency distribution but does not explicitly account for co-tagging correlations. Even under these constraints, we demonstrate empirically and theoretically that our model can reproduce a number of statistical properties of the co-tagging graph that links tags appearing in the same post.