2021/08/06 by Kenny Peng, Arunesh Mathur, Peng, Kenny +3 · 1 voice · 8 citations
Computer Science · Social Sciences · #CLARITY #Computer science #Computer security #Data science #Engineering ethics #Ethical issues #Ethics and Social Impacts of AI #Face (sociological concept) #Face recognition and analysis #Political science #Premise #Privacy-Preserving Technologies in Data #Process (computing) #Sociology #Stewardship (theology) #Transparency (behavior) #cs.CY #cs.LG
paper · pdf · doi:10.48550/arxiv.2108.02922
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2021/08/06 · arxiv created 2021/11/21 · arxiv updated 2021/11/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Machine learning datasets have elicited concerns about privacy, bias, and unethical applications, leading to the retraction of prominent datasets such as DukeMTMC, MS-Celeb-1M, and Tiny Images. In response, the machine learning community has called for higher ethical standards in dataset creation. To help inform these efforts, we studied three influential but ethically problematic face and person recognition datasets -- Labeled Faces in the Wild (LFW), MS-Celeb-1M, and DukeMTM -- by analyzing nearly 1000 papers that cite them. We found that the creation of derivative datasets and models, broader technological and social change, the lack of clarity of licenses, and dataset management practices can introduce a wide range of ethical concerns. We conclude by suggesting a distributed approach to harm mitigation that considers the entire life cycle of a dataset.