vix.ing · top · new · best · stats · spec

Discourse-aware rumour stance classification in social media using sequential classifiers

2017/12/06 by Arkaitz Zubiaga, Elena Kochkina, Maria Liakata +6
Computer Science · Mathematics · Physics and Astronomy · Social Sciences · #Artificial intelligence #Classifier (UML) #Complex Network Analysis Techniques #Computer science #Conditional random field #Decision tree #Exploit #Machine learning #Mathematics #Misinformation and Its Impacts #Natural language processing #Opinion Dynamics and Social Influence #Random forest #Random subspace method #Set (abstract data type) #Social media #Tree (set theory) #World Wide Web #cs.CL #cs.SI

paper · pdf · doi:10.1016/j.ipm.2017.11.009

published as Information Processing & Management, Volume 54, Issue 2, March 2018, Pages 273-290

arxiv created 2017/12/06 · openalex publication_date 2017/12/06 · arxiv updated 2017/12/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Rumour stance classification, defined as classifying the stance of specific social media posts into one of supporting, denying, querying or commenting on an earlier post, is becoming of increasing interest to researchers. While most previous work has focused on using individual tweets as classifier inputs, here we report on the performance of sequential classifiers that exploit the discourse features inherent in social media interactions or 'conversational threads'. Testing the effectiveness of four sequential classifiers -- Hawkes Processes, Linear-Chain Conditional Random Fields (Linear CRF), Tree-Structured Conditional Random Fields (Tree CRF) and Long Short Term Memory networks (LSTM) -- on eight datasets associated with breaking news stories, and looking at different types of local and contextual features, our work sheds new light on the development of accurate stance classifiers. We show that sequential classifiers that exploit the use of discourse properties in social media conversations while using only local features, outperform non-sequential classifiers. Furthermore, we show that LSTM using a reduced set of features can outperform the other sequential classifiers; this performance is consistent across datasets and across types of stances. To conclude, our work also analyses the different features under study, identifying those that best help characterise and distinguish between stances, such as supporting tweets being more likely to be accompanied by evidence than denying tweets. We also set forth a number of directions for future research.

Citations