2023/09/20 by Yida Mu, Xingyi Song, Mu, Yida +5 · 1 citation
Computer Science · Physics and Astronomy · Social Sciences · #Complex Network Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Misinformation and Its Impacts #Spam and Phishing Detection
paper · pdf · doi:10.48550/arxiv.2309.11576
openalex publication_date 2023/09/20 · openalex created_date 2023/09/23 · openalex updated_date 2026/07/28
A crucial aspect of a rumor detection model is its ability to generalize, particularly its ability to detect emerging, previously unknown rumors. Past research has indicated that content-based (i.e., using solely source posts as input) rumor detection models tend to perform less effectively on unseen rumors. At the same time, the potential of context-based models remains largely untapped. The main contribution of this paper is in the in-depth evaluation of the performance gap between content and context-based models specifically on detecting new, unseen rumors. Our empirical findings demonstrate that context-based models are still overly dependent on the information derived from the rumors' source post and tend to overlook the significant role that contextual information can play. We also study the effect of data split strategies on classifier performance. Based on our experimental results, the paper also offers practical suggestions on how to minimize the effects of temporal concept drift in static datasets during the training of rumor detection methods.