2020/02/14 by Irene Kyomuhangi, Tarekegn A. Abeku, Kyomuhangi, Irene +7
Biochemistry, Genetics and Molecular Biology · #Applications (stat.AP) #FOS: Computer and information sciences #Genetic Associations and Epidemiology #Methodology (stat.ME)
paper · pdf · doi:10.48550/arxiv.2002.06032
openalex publication_date 2020/02/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Diagnosis is often based on the exceedance or not of continuous health\nindicators of a predefined cut-off value, so as to classify patients into\npositives and negatives for the disease under investigation. In this paper, we\ninvestigate the effects of dichotomization of spatially-referenced continuous\noutcome variables on geostatistical inference. Although this issue has been\nextensively studied in other fields, dichotomization is still a common practice\nin epidemiological studies. Furthermore, the effects of this practice in the\ncontext of prevalence mapping have not been fully understood. Here, we\ndemonstrate how spatial correlation affects the loss of information due to\ndichotomization, how linear geostatistical models can be used to map disease\nprevalence and thus avoid dichotomization, and finally, how dichotomization\naffects our predictive inference on prevalence. To pursue these objectives, we\ndevelop a metric, based on the composite likelihood, which can be used to\nquantify the potential loss of information after dichotomization without\nrequiring the fitting of Binomial geostatistical models. Through a simulation\nstudy and two applications on disease mapping in Africa, we show that, as\nthresholds used for dichotomization move further away from the mean of the\nunderlying process, the performance of binomial geostatistical models\ndeteriorates substantially. We also find that dichotomization can lead to the\nloss of fine scale features of disease prevalence and increased uncertainty in\nthe parameter estimates, especially in the presence of a large noise to signal\nratio. These findings strongly support the conclusions from previous studies\nthat dichotomization should be always avoided whenever feasible.\n