2026/01/31 by Danielle Tsao, Krikamol Muandet, Frederick Eberhardt +1
Mathematics · #stat.ME #stat.ML
arxiv created 2026/08/01 · arxiv updated 2026/08/04
Instrumental variable estimation has emerged as a standard approach to mitigating confounding bias in the social sciences and epidemiology, where conducting randomized experiments can be too costly or infeasible. However, justifying the validity of the instrument is frequently challenging. We highlight a problem generally neglected in arguments for instrumental variable validity: the presence of an "aggregate treatment variable", where the treatment (e.g., education, GDP, caloric intake) is composed of finer-grained, unobserved components that each may have a different effect on the outcome. While the aggregation problem itself is general, our focus is on instrumental variable estimation in a linear setting, the regime underlying much of applied IV practice. We show that the causal effect of an aggregate treatment is generally ambiguous, as it depends on how an intervention on the aggregate is instantiated at the component level. We formalize this relation using the aggregate-constrained component intervention distribution (ACID). We then identify two key conditions under which standard instrumental variable estimators identify the aggregate effect. The contrived nature of these conditions implies major limitations on the interpretation of instrumental variable estimates based on aggregate treatments and highlights the need for a broader justificatory base for the exclusion restriction in such settings.