vix.ing · top · new · best · stats · spec

Semantics derived automatically from language corpora contain human-like biases

2016/08/31 by Aylin Caliskan, Joanna J. Bryson, Arvind Narayanan · 5 citations
Computer Science · Neuroscience · Psychology · Social Sciences · #Artificial intelligence #Authorship Attribution and Profiling #Cognitive psychology #Computer science #Embedding #Ethics and Social Impacts of AI #Information retrieval #Linguistics #Natural language processing #Prejudice (legal term) #Psychology #Psychology of Moral and Emotional Judgment #Replicate #Semantics (computer science) #Social psychology #Test (biology) #Word (group theory) #Word embedding #cs.AI #cs.CL #cs.CY #cs.LG #sort

paper · pdf · doi:10.1126/science.aal4230

14 pages, 3 figures

openalex publication_date 2017/04/13 · arxiv created 2017/05/25 · arxiv updated 2017/05/26 · openalex created_date 2020/11/23 · openalex updated_date 2026/08/05

Abstract

Machine learning is a means to derive artificial intelligence by discovering patterns in existing data. Here, we show that applying machine learning to ordinary human language results in human-like semantic biases. We replicated a spectrum of known biases, as measured by the Implicit Association Test, using a widely used, purely statistical machine-learning model trained on a standard corpus of text from the World Wide Web. Our results indicate that text corpora contain recoverable and accurate imprints of our historic biases, whether morally neutral as toward insects or flowers, problematic as toward race or gender, or even simply veridical, reflecting the status quo distribution of gender with respect to careers or first names. Our methods hold promise for identifying and addressing sources of bias in culture, including technology.

Citations

Cited by