2020/05/07 by Rohan Arambepola, Tim Lucas, Arambepola, Rohan +7
Biochemistry, Genetics and Molecular Biology · Economics, Econometrics and Finance · #Applications (stat.AP) #FOS: Computer and information sciences #Genetic and phenotypic traits in livestock #Health Systems, Economic Evaluations, Quality of Life #Spatial and Panel Data Analysis
paper · pdf · doi:10.48550/arxiv.2005.03604
openalex publication_date 2020/05/07 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
Disaggregation regression has become an important tool in spatial disease\nmapping for making fine-scale predictions of disease risk from aggregated\nresponse data. By including high resolution covariate information and modelling\nthe data generating process on a fine scale, it is hoped that these models can\naccurately learn the relationships between covariates and response at a fine\nspatial scale. However, validating these high resolution predictions can be a\nchallenge, as often there is no data observed at this spatial scale. In this\nstudy, disaggregation regression was performed on simulated data in various\nsettings and the resulting fine-scale predictions are compared to the simulated\nground truth. Performance was investigated with varying numbers of data points,\nsizes of aggregated areas and levels of model misspecification. The\neffectiveness of cross validation on the aggregate level as a measure of\nfine-scale predictive performance was also investigated. Predictive performance\nimproved as the number of observations increased and as the size of the\naggregated areas decreased. When the model was well-specified, fine-scale\npredictions were accurate even with small numbers of observations and large\naggregated areas. Under model misspecification predictive performance was\nsignificantly worse for large aggregated areas but remained high when response\ndata was aggregated over smaller regions. Cross-validation correlation on the\naggregate level was a moderately good predictor of fine-scale predictive\nperformance. While the simulations are unlikely to capture the nuances of\nreal-life response data, this study gives insight into the effectiveness of\ndisaggregation regression in different contexts.\n