2021/06/06 by Jaime Roquero Gimenez, Dominik Rothenhaüsler, Gimenez, Jaime Roquero +1
Computer Science · Mathematics · #Bayesian Modeling and Causal Inference #Advanced Causal Inference Techniques #Statistical Methods and Inference
paper · pdf · doi:10.48550/arxiv.2106.03024
In causal inference, it is common to estimate the causal effect of a single\ntreatment variable on an outcome. However, practitioners may also be interested\nin the effect of simultaneous interventions on multiple covariates of a fixed\ntarget variable. We propose a novel method that allows to estimate the effect\nof joint interventions using data from different experiments in which only very\nfew variables are manipulated. If there is only little randomized data or no\nrandomized data at all, one can use observational data sets if certain parental\nsets are known or instrumental variables are available. If the joint causal\neffect is linear, the proposed method can be used for estimation and inference\nof joint causal effects, and we characterize conditions for identifiability. In\nthe overidentified case, we indicate how to leverage all the available causal\ninformation across multiple data sets to efficiently estimate the causal\neffects. If the dimension of the covariate vector is large, we may only have a\nfew samples in each data set. Under a sparsity assumption, we derive an\nestimator of the causal effects in this high-dimensional scenario. In addition,\nwe show how to deal with the case where a lack of experimental constraints\nprevents direct estimation of the causal effects. When the joint causal effects\nare non-linear, we characterize conditions under which identifiability holds,\nand propose a non-linear causal aggregation methodology for experimental data\nsets similar to the gradient boosting algorithm where in each iteration we\ncombine weak learners trained on different datasets using only unconfounded\nsamples. We demonstrate the effectiveness of the proposed method on simulated\nand semi-synthetic data.\n