Can NLP Tackle Hate Speech in the Real World? Stakeholder-Informed Feedback and Survey on Counterspeech
2025/08/06 by Dinkar, Tanvi, Jiang, Aiqi, Frenda, Simona +4
#Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2508.04638
Abstract
Counterspeech, i.e. the practice of responding to online hate speech, has gained traction in NLP as a promising intervention. While early work emphasised collaboration with non-governmental organisation stakeholders, recent research trends have shifted toward automated pipelines that reuse a small set of legacy datasets, often without input from affected communities. This paper presents a systematic review of 74 NLP studies on counterspeech, analysing the extent to which stakeholder participation influences dataset creation, model development, and evaluation. To complement this analysis, we conducted a participatory case study with five NGOs specialising in online Gender-Based Violence (oGBV), identifying stakeholder-informed practices for counterspeech generation. Our findings reveal a growing disconnect between current NLP research and the needs of communities most impacted by toxic online content. We conclude with concrete recommendations for re-centring stakeholder expertise in counterspeech research.
Citations
- CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs
- Echoes of Discord: Forecasting Hater Reactions to Counterspeech
- PANDA -- Paired Anti-hate Narratives Dataset from Asia: Using an LLM-as-a-Judge to Create the First Chinese Counterspeech Dataset
- Generative AI may backfire for counterspeech
- Perceiving and Countering Hate: The Role of Identity in Online Responses
- Rescuing Counterspeech: A Bridging-Based Approach to Combating Misinformation
- Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering
- CrowdCounter: A benchmark type-specific multi-target counterspeech dataset
- COT: A Generative Approach for Hate Speech Counter-Narratives via Contrastive Optimal Transport
- Hostile Counterspeech Drives Users From Hate Subreddits
- NLP Systems That Can't Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps
- NLP for Counterspeech against Hate: A Survey and How-To Guide
- Gendered Inequalities in Online Harms: Fear, Safety Work, and Online Participation
- Hatred Stems from Ignorance! Distillation of the Persuasion Modes in Countering Conversational Hate Speech
- Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF
- Basque and Spanish Counter Narrative Generation: Data Creation and Evaluation
- Counterspeakers' Perspectives: Unveiling Barriers and AI Needs in the Fight against Online Hate
- Low-Resource Counterspeech Generation for Indic Languages: The Case of Bengali and Hindi
- Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
- Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
- Alternative Speech: Complementary Method to Counter-Narrative for Better Discourse
- ReZG: Retrieval-Augmented Zero-Shot Counter Narrative Generation for Hate Speech
- The Participatory Turn in AI Design: Theoretical Foundations and the Current State of Practice
- Weigh Your Own Words: Improving Hate Speech Counter Narrative Generation via Attention Regularization
- Understanding Counterspeech for Online Harm Mitigation
- Which Argumentative Aspects of Hate Speech in Social Media can be reliably identified?
- Counterspeeches up my sleeve! Intent Distribution Learning and Persistent Fusion for Intent-Conditioned Counterspeech Generation
- Reinforcement Learning-based Counter-Misinformation Response Generation: A Case Study of COVID-19 Vaccine Misinformation
- SemEval-2023 Task 10: Explainable Detection of Online Sexism
- Human-Machine Collaboration Approaches to Build a Dialogue Dataset for Hate Speech Countering
- Birdwatch: Crowd Wisdom and Bridging Algorithms can Inform Understanding and Reduce the Spread of Misinformation
- Power to the People? Opportunities and Challenges for Participatory AI
- KOLD: Korean Offensive Language Dataset
- CounterGeDi: A controllable approach to generate polite, detoxified and emotional counterspeech
- Using Pre-Trained Language Models for Producing Counter Narratives Against Hate Speech: a Comparative Study
- APEACH: Attacking Pejorative Expressions with Analysis on Crowd-Generated Hate Speech Evaluation Datasets
- Counter Hate Speech in Social Media: A Survey
- COLD: A Benchmark for Chinese Offensive Language Detection
- Multilingual Counter Narrative Type Classification
- SWSR: A Chinese Dataset and Lexicon for Online Sexism Detection
- Towards Knowledge-Grounded Counter Narrative Generation for Hate Speech
- Generate, Prune, Select: A Pipeline for Counterspeech Generation against Online Hate Speech
- HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection
- Countering hate on social media: Large scale classification of hate and counter speech
- BEEP! Korean Corpus of Online News Comments for Toxic Speech Detection
- Racism is a Virus: Anti-Asian Hate and Counterspeech in Social Media during the COVID-19 Crisis
- Generating Counter Narratives against Online Hate Speech: Data and\n Strategies
- Social Bias Frames: Reasoning about Social and Power Implications of\n Language
- Multi-label Categorization of Accounts of Sexism using a Neural\n Framework
- A Benchmark Dataset for Learning to Intervene in Online Hate Speech
- Multilingual and Multi-Aspect Hate Speech Analysis
- Analyzing the hate and counter speech accounts on Twitter
- Hate Speech Dataset from a White Supremacy Forum
- Designing Human-AI Collaboration to Support Learning in Counterspeech Writing
Related