vix.ing · top · new · best · stats · spec

Improved and Robust Controversy Detection in General Web Pages Using\n Semantic Approaches under Large Scale Conditions

2018/12/02 by Jasper Linmans, Linmans, Jasper, Bob van de Velde +3
Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Software Engineering Research #Topic Modeling #Wikis in Education and Collaboration

paper · pdf · doi:10.48550/arxiv.1812.00382

openalex publication_date 2018/12/02 · openalex created_date 2022/08/01 · openalex updated_date 2026/07/28

Abstract

Detecting controversy in general web pages is a daunting task, but\nincreasingly essential to efficiently moderate discussions and effectively\nfilter problematic content. Unfortunately, controversies occur across many\ntopics and domains, with great changes over time. This paper investigates\nneural classifiers as a more robust methodology for controversy detection in\ngeneral web pages. Current models have often cast controversy detection on\ngeneral web pages as Wikipedia linking, or exact lexical matching tasks. The\ndiverse and changing nature of controversies suggest that semantic approaches\nare better able to detect controversy. We train neural networks that can\ncapture semantic information from texts using weak signal data. By leveraging\nthe semantic properties of word embeddings we robustly improve on existing\ncontroversy detection methods. To evaluate model stability over time and to\nunseen topics, we asses model performance under varying training conditions to\ntest cross-temporal, cross-topic, cross-domain performance and annotator\ncongruence. In doing so, we demonstrate that weak-signal based neural\napproaches are closer to human estimates of controversy and are more robust to\nthe inherent variability of controversies.\n

Citations

Related