2017/08/23 by Yilin Zhang, Zhang, Yilin, Marie Poux-Berthe +8 · 3 citations
Computer Science · Mathematics · Physics and Astronomy · Social Sciences · #Applications (stat.AP) #Artificial intelligence #Cluster analysis #Complex Network Analysis Techniques #Computation and Language (cs.CL) #Computer science #Contextualization #Data science #FOS: Computer and information sciences #FOS: Physical sciences #Graph #Information retrieval #Opinion Dynamics and Social Influence #Physics and Society (physics.soc-ph) #Political science #Politics #Presidential election #Social Media and Politics #Theoretical computer science #cs.CL #physics.soc-ph #stat.AP
paper · pdf · doi:10.48550/arxiv.1708.06872
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2017/08/23 · arxiv created 2018/03/24 · arxiv updated 2018/03/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We propose a graph contextualization method, pairGraphText, to study political engagement on Facebook during the 2012 French presidential election. It is a spectral algorithm that contextualizes graph data with text data for online discussion thread. In particular, we examine the Facebook posts of the eight leading candidates and the comments beneath these posts. We find evidence of both (i) candidate-centered structure, where citizens primarily comment on the wall of one candidate and (ii) issue-centered structure (i.e. on political topics), where citizens' attention and expression is primarily directed towards a specific set of issues (e.g. economics, immigration, etc). To identify issue-centered structure, we develop pairGraphText, to analyze a network with high-dimensional features on the interactions (i.e. text). This technique scales to hundreds of thousands of nodes and thousands of unique words. In the Facebook data, spectral clustering without the contextualizing text information finds a mixture of (i) candidate and (ii) issue clusters. The contextualized information with text data helps to separate these two structures. We conclude by showing that the novel methodology is consistent under a statistical model.