vix.ing · top · new · best · stats

A statistical framework for fair predictive algorithms

2016/10/25 by Kristian Lum, Lum, Kristian, James E. Johndrow +2 · 58 citations
Computer Science · Mathematics · Psychology · Social Sciences · #Adversarial Robustness in Machine Learning #Algorithm #Artificial intelligence #Binary number #Computer science #Covariate #Criminal Justice and Corrections Analysis #Criminal justice #Criminology #Disparate impact #Disparate treatment #Ethics and Social Impacts of AI #Law #Machine learning #Mathematics #Neutrality #Political science #Prejudice (legal term) #Probabilistic logic #Process (computing) #Psychology #Race (biology) #Relevance (law) #Social psychology #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.1610.08077

published in arXiv (Cornell University) (Cornell University)

arxiv created 2016/10/25 · openalex publication_date 2016/10/25 · arxiv updated 2016/10/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Predictive modeling is increasingly being employed to assist human decision-makers. One purported advantage of replacing human judgment with computer models in high stakes settings-- such as sentencing, hiring, policing, college admissions, and parole decisions-- is the perceived "neutrality" of computers. It is argued that because computer models do not hold personal prejudice, the predictions they produce will be equally free from prejudice. There is growing recognition that employing algorithms does not remove the potential for bias, and can even amplify it, since training data were inevitably generated by a process that is itself biased. In this paper, we provide a probabilistic definition of algorithmic bias. We propose a method to remove bias from predictive models by removing all information regarding protected variables from the permitted training data. Unlike previous work in this area, our framework is general enough to accommodate arbitrary data types, e.g. binary, continuous, etc. Motivated by models currently in use in the criminal justice system that inform decisions on pre-trial release and paroling, we apply our proposed method to a dataset on the criminal histories of individuals at the time of sentencing to produce "race-neutral" predictions of re-arrest. In the process, we demonstrate that the most common approach to creating "race-neutral" models-- omitting race as a covariate-- still results in racially disparate predictions. We then demonstrate that the application of our proposed method to these data removes racial disparities from predictions with minimal impact on predictive accuracy.

Cited by

Related