StereoSet: Measuring stereotypical bias in pretrained language models
2020/04/20 by Nadeem, Moin, Bethke, Anna, Reddy, Siva · 93 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2004.09456
Abstract
A stereotype is an over-generalized belief about a particular group of people, e.g., Asians are good at math or Asians are bad drivers. Such beliefs (biases) are known to hurt target groups. Since pretrained language models are trained on large real world data, they are known to capture stereotypical biases. In order to assess the adverse effects of these models, it is important to quantify the bias captured in them. Existing literature on quantifying bias evaluates pretrained language models on a small set of artificially constructed bias-assessing sentences. We present StereoSet, a large-scale natural dataset in English to measure stereotypical biases in four domains: gender, profession, race, and religion. We evaluate popular models like BERT, GPT-2, RoBERTa, and XLNet on our dataset and show that these models exhibit strong stereotypical biases. We also present a leaderboard with a hidden test set to track the bias of future language models at https://stereoset.mit.edu
Cited by
- Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context
- Responsible Intelligence in Practice: A Fairness Audit of Open Large Language Models for Library Reference Services
- Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
- AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models
- An Empirical Survey of Model Merging Algorithms for Social Bias Mitigation
- Understanding Down Syndrome Stereotypes in LLM-Based Personas
- BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models
- Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
- MoodBench 1.0: An Evaluation Benchmark for Emotional Companionship Dialogue Systems
- Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
- Annotating Dimensions of Social Perception in Text: A Sentence-Level Dataset of Warmth and Competence
- HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detection
- Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
- Large Language Models Develop Novel Social Biases Through Adaptive Exploration
- Multi-Reward GRPO Fine-Tuning for De-biasing Large Language Models: A Study Based on Chinese-Context Discrimination Data
- Silenced Biases: The Dark Side LLMs Learned to Refuse
- Surfacing Subtle Stereotypes: A Multilingual, Debate-Oriented Evaluation of Modern LLMs
- Controlling Gender Bias in Retrieval via a Backpack Architecture
- TriCon-Fair: Triplet Contrastive Learning for Mitigating Social Bias in Pre-trained Language Models
- Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
- Biases in the Blind Spot: Detecting What LLMs Fail to Mention
- Adaptive Data Collection for Latin-American Community-sourced Evaluation of Stereotypes (LACES)
- A word association network methodology for evaluating implicit biases in LLMs compared to humans
- Breaking the Benchmark: Revealing LLM Bias via Minimal Contextual Augmentation
- PBBQ: A Persian Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models
- Investigating Thinking Behaviours of Reasoning-Based Language Models for Social Bias Mitigation
- Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
- HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
- The Social Cost of Intelligence: Emergence, Propagation, and Amplification of Stereotypical Bias in Multi-Agent Systems
- SAGE: A Top-Down Bottom-Up Knowledge-Grounded User Simulator for Multi-turn AGent Evaluation
- ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
- CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs
- MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
- Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions
- Textual Entailment and Token Probability as Bias Evaluation Metrics
- Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups
- LLM Bias Detection and Mitigation through the Lens of Desired Distributions
- EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
- Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
- Homophily-induced Emergence of Biased Structures in LLM-based Multi-Agent AI Systems
- When Voice Matters: Evidence of Gender Disparity in Positional Bias of SpeechLLMs
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
- RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
- Mitigating Biases in Language Models via Bias Unlearning
- Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Dataset
- Act as an expert in psychometry. The evaluation of large language models utility in psychological tests cross-cultural adaptations
- BTC-SAM: Leveraging LLMs for Generation of Bias Test Cases for Sentiment Analysis Models
- GeoBS: Information-Theoretic Quantification of Geographic Bias in AI Models
- Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
- Language, Culture, and Ideology: Personalizing Offensiveness Detection in Political Tweets with Reasoning LLMs
- Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias
- Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
- Acoustic-based Gender Differentiation in Speech-aware Language Models
- Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
- SMITE: Enhancing Fairness in LLMs through Optimal In-Context Example Selection via Dynamic Validation
- AccessEval: Benchmarking Disability Bias in Large Language Models
- Mechanistic Interpretability with SAEs: Probing Religion, Violence, and Geography in Large Language Models
- nDNA -- the Semantic Helix of Artificial Cognition
- Intrinsic Meets Extrinsic Fairness: Assessing the Downstream Impact of Bias Mitigation in Large Language Models
- Fair-GPTQ: Bias-Aware Quantization for Large Language Models
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- Simulating a Bias Mitigation Scenario in Large Language Models
- Stochastic Streets: A Walk Through Random LLM Address Generation in four European Cities
- Don't Change My View: Ideological Bias Auditing in Large Language Models
- MetaRAG: Metamorphic Testing for Hallucination Detection in RAG Systems
- Simulating Identity, Propagating Bias: Abstraction and Stereotypes in LLM-Generated Text
- Bias after Prompting: Persistent Discrimination in Large Language Models
- Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases
- Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making
- Revealing Potential Biases in LLM-Based Recommender Systems in the Cold Start Setting
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- How Quantization Shapes Bias in Large Language Models
- Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience
- Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
- Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
- DAIQ: Auditing Demographic Attribute Inference from Question in LLMs
- Spot the BlindSpots: Systematic Identification and Quantification of Fine-Grained LLM Biases in Contact Center Summaries
- The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases
- Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
- BIPOLAR: Polarization-based granular framework for LLM bias evaluation
- BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
- Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
- BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context
- Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
- Semantic and Structural Analysis of Implicit Biases in Large Language Models: An Interpretable Approach
- Do Biased Models Have Biased Thoughts?
- The World According to LLMs: How Geographic Origin Influences LLMs' Entity Deduction Capabilities
- I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
- FairLangProc: A Python package for fairness in NLP
- Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes
- Bias Association Discovery Framework for Open-Ended LLM Generations
- A Transparent Fairness Evaluation Protocol for Open-Source Language Model Benchmarking on the Blockchain
Related