2024/07/22 by Krish Muralidhar, Steven Ruggles, Muralidhar, Krish +1
Mathematics · Social Sciences · #Census and Population Estimation #Databases (cs.DB) #FOS: Computer and information sciences #Gender, Labor, and Family Dynamics #Health disparities and outcomes
paper · pdf · doi:10.48550/arxiv.2407.15957
openalex publication_date 2024/07/22 · openalex created_date 2025/01/05 · openalex updated_date 2026/07/28
In 2017, the United States Census Bureau announced that because of high disclosure risk in the methodology (data swapping) used to produce tabular data for the 2010 census, a different protection mechanism based on differential privacy would be used for the 2020 census. While there have been many studies evaluating the result of this change, there has been no rigorous examination of disclosure risk claims resulting from the released 2010 tabular data. In this study we perform such an evaluation. We show that the procedures used to evaluate disclosure risk are unreliable and resulted in inflated disclosure risk. Demonstration data products released using the new procedure were also shown to have poor utility. However, since the Census Bureau had already committed to a different procedure, they had no option except to escalate their commitment. The result of such escalation is that the 2020 tabular data release offers neither privacy nor accuracy.