vix.ing · top · new · best · stats · spec

Explainability and Hate Speech: Structured Explanations Make Social Media Moderators Faster

2024/06/06 by Agostina Calabrese, Calabrese, Agostina, Leonardo Neves +11 · 3 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection

paper · pdf · doi:10.48550/arxiv.2406.04106

openalex publication_date 2024/06/06 · openalex created_date 2024/06/08 · openalex updated_date 2026/07/28

Abstract

Content moderators play a key role in keeping the conversation on social media healthy. While the high volume of content they need to judge represents a bottleneck to the moderation pipeline, no studies have explored how models could support them to make faster decisions. There is, by now, a vast body of research into detecting hate speech, sometimes explicitly motivated by a desire to help improve content moderation, but published research using real content moderators is scarce. In this work we investigate the effect of explanations on the speed of real-world moderators. Our experiments show that while generic explanations do not affect their speed and are often ignored, structured explanations lower moderators' decision making time by 7.4%.

Cited by

Related