2021/04/16 by Xiaonan Jing, Qingyuan Hu, Jing, Xiaonan +5
Computer Science · Physics and Astronomy · #Complex Network Analysis Techniques #Computation and Language (cs.CL) #Data Management and Algorithms #FOS: Computer and information sciences #Opinion Dynamics and Social Influence #cs.CL
paper · pdf · doi:10.48550/arxiv.2104.07836
Accepted as full paper by the 34th International FLAIRS Conference
arxiv created 2021/04/16 · openalex publication_date 2021/04/16 · arxiv updated 2021/04/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Twitter serves as a data source for many Natural Language Processing (NLP) tasks. It can be challenging to identify topics on Twitter due to continuous updating data stream. In this paper, we present an unsupervised graph based framework to identify the evolution of sub-topics within two weeks of real-world Twitter data. We first employ a Markov Clustering Algorithm (MCL) with a node removal method to identify optimal graph clusters from temporal Graph-of-Words (GoW). Subsequently, we model the clustering transitions between the temporal graphs to identify the topic evolution. Finally, the transition flows generated from both computational approach and human annotations are compared to ensure the validity of our framework.