2019/01/29 by Bokun Wang, Wang, Bokun, Ian Davidson +1 · 2 citations
Computer Science · Medicine · #Data-Driven Disease Surveillance #FOS: Computer and information sciences #HIV, Drug Use, Sexual Risk #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Privacy-Preserving Technologies in Data
paper · pdf · doi:10.48550/arxiv.1901.10053
openalex publication_date 2019/01/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Fair clustering under the disparate impact doctrine requires that population of each protected group should be approximately equal in every cluster. Previous work investigated a difficult-to-scale pre-processing step for k-center and k-median style algorithms for the special case of this problem when the number of protected groups is two. In this work, we consider a more general and practical setting where there can be many protected groups. To this end, we propose Deep Fair Clustering, which learns a discriminative but fair cluster assignment function. The experimental results on three public datasets with different types of protected attribute show that our approach can steadily improve the degree of fairness while only having minor loss in terms of clustering quality.