2025/10/06 by Feihao Chen, Majumdar, Ayan, Jinghui Li +4 · 1 citation
Computer Science · #Audit #Benchmark (surveying) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Focus (optics) #Machine Learning (cs.LG) #Regulatory focus theory #Scalability #Sentiment Analysis and Opinion Mining #Social media #Strengths and weaknesses #Task (project management)
paper · pdf · doi:10.48550/arxiv.2510.04641
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/10/06 · openalex created_date 2025/10/09 · openalex updated_date 2026/07/28
Large-scale web-scraped text corpora used to train general-purpose AI models often contain harmful demographic-targeted social biases, creating a regulatory need for data auditing and developing scalable bias-detection methods. Although prior work has investigated biases in text datasets and related detection methods, these studies remain narrow in scope. They typically focus on a single content type (e.g., hate speech), cover limited demographic axes, overlook biases affecting multiple demographics simultaneously, and analyze limited techniques. Consequently, practitioners lack a holistic understanding of the strengths and limitations of recent large language models (LLMs) for automated bias detection. In this study, we conduct a comprehensive benchmark study on English texts to assess the ability of LLMs in detecting demographic-targeted social biases. To align with regulatory requirements, we frame bias detection as a multi-label task of detecting targeted identities using a demographic-focused taxonomy. We then systematically evaluate models across scales and techniques, including prompting, in-context learning, and fine-tuning. Using twelve datasets spanning diverse content types and demographics, our study demonstrates the promise of fine-tuned smaller models for scalable detection. However, our analyses also expose persistent gaps across demographic axes and multi-demographic targeted biases, underscoring the need for more effective and scalable detection frameworks.