vix.ing · top · new · best · stats · spec

A causal inference framework for cancer cluster investigations using\n publicly available data

2018/11/14 by Rachel C. Nethery, Nethery, Rachel C., Yue Yang +5
Mathematics · Medicine · #Advanced Causal Inference Techniques #Applications (stat.AP) #Data-Driven Disease Surveillance #FOS: Computer and information sciences #Methodology (stat.ME) #Statistical Methods and Bayesian Inference

paper · pdf · doi:10.48550/arxiv.1811.05997

openalex publication_date 2018/11/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Often, a community becomes alarmed when high rates of cancer are noticed, and\nresidents suspect that the cancer cases could be caused by a known source of\nhazard. In response, the CDC recommends that departments of health perform a\nstandardized incidence ratio (SIR) analysis to determine whether the observed\ncancer incidence is higher than expected. This approach has several limitations\nthat are well documented in the literature. In this paper we propose a novel\ncausal inference approach to cancer cluster investigations, rooted in the\npotential outcomes framework. Assuming that a source of hazard representing a\npotential cause of increased cancer rates in the community is identified a\npriori, we introduce a new estimand called the causal SIR (cSIR). The cSIR is a\nratio defined as the expected cancer incidence in the exposed population\ndivided by the expected cancer incidence under the (counterfactual) scenario of\nno exposure. To estimate the cSIR we need to overcome two main challenges: 1)\nidentify unexposed populations that are as similar as possible to the exposed\none to inform estimation under the counterfactual scenario of no exposure, and\n2) make inference on cancer incidence in these unexposed populations using\npublicly available data that are often available at a much higher level of\nspatial aggregation than what is desired. We overcome the first challenge by\nrelying on matching. We overcome the second challenge by developing a Bayesian\nhierarchical model that borrows information from other sources to impute cancer\nincidence at the desired finer level of spatial aggregation. We apply our\nproposed approach to determine whether trichloroethylene vapor exposure has\ncaused increased cancer incidence in Endicott, NY.\n

Related