vix.ing · top · new · best · stats · spec

Semi-supervised Content-based Detection of Misinformation via Tensor\n Embeddings

2018/04/24 by Gisel Bastidas Guacho, Guacho, Gisel Bastidas, Sara Abdali +5 · 2 citations
Computer Science · Social Sciences · #Applications (stat.AP) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Misinformation and Its Impacts #Social and Information Networks (cs.SI) #Spam and Phishing Detection #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1804.09088

openalex publication_date 2018/04/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Fake news may be intentionally created to promote economic, political and\nsocial interests, and can lead to negative impacts on humans beliefs and\ndecisions. Hence, detection of fake news is an emerging problem that has become\nextremely prevalent during the last few years. Most existing works on this\ntopic focus on manual feature extraction and supervised classification models\nleveraging a large number of labeled (fake or real) articles. In contrast, we\nfocus on content-based detection of fake news articles, while assuming that we\nhave a small amount of labels, made available by manual fact-checkers or\nautomated sources. We argue this is a more realistic setting in the presence of\nmassive amounts of content, most of which cannot be easily factchecked. To that\nend, we represent collections of news articles as multi-dimensional tensors,\nleverage tensor decomposition to derive concise article embeddings that capture\nspatial/contextual information about each news article, and use those\nembeddings to create an article-by-article graph on which we propagate limited\nlabels. Results on three real-world datasets show that our method performs on\npar or better than existing models that are fully supervised, in that we\nachieve better detection accuracy using fewer labels. In particular, our\nproposed method achieves 75.43% of accuracy using only 30% of labels of a\npublic dataset while an SVM-based classifier achieved 67.43%. Furthermore, our\nmethod achieves 70.92% of accuracy in a large dataset using only 2% of labels.\n

Cited by

Related