2013/06/20 by Gustavo Pabón, Claudio Gutiérrez, Javier D. Fernández +1 · 1 citation
Computer Science · Decision Sciences · #Semantic Web and Ontologies #Data Quality and Management #Microdata (statistics) #Census #Publication #Interoperability #Computer science #American Community Survey #Data science #Population #Data mining #Geography #Information retrieval #World Wide Web #Database #Sociology #Demography #Political science
paper · pdf · doi:10.1002/asi.22876
openalex publication_date 2013/06/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Censuses are one of the most relevant types of statistical data, allowing analyses of the population in terms of demography, economy, sociology, and culture. For fine‐grained analysis, census agencies publish census microdata that consist of a sample of individual records of the census containing detailed anonymous individual information. Working with microdata from different censuses and doing comparative studies are currently difficult tasks due to the diversity of formats and granularities. In this article, we show that novel data processing techniques can be applied to make census microdata interoperable and easy to access and combine. In fact, we demonstrate how L inked O pen D ata principles, a set of techniques to publish and make connections of (semi‐)structured data on the web, can be fruitfully applied to census microdata. We present a step‐by‐step process to achieve this goal and we study, in theory and practice, two real case studies: the 2001 Spanish census and a general framework for I ntegrated P ublic U se M icrodata S eries ( IPUMS ‐I).