2024/03/21 by Jonathan Fuhr, Philipp Berens, Fuhr, Jonathan +3 · 1 citation
Computer Science · #Econometrics (econ.EM) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #FOS: Economics and business #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification #Methodology (stat.ME)
paper · pdf · doi:10.48550/arxiv.2403.14385
openalex publication_date 2024/03/21 · openalex created_date 2024/03/24 · openalex updated_date 2026/08/01
The estimation of causal effects with observational data continues to be a very active research area. In recent years, researchers have developed new frameworks which use machine learning to relax classical assumptions necessary for the estimation of causal effects. In this paper, we review one of the most prominent methods - "double/debiased machine learning" (DML) - and empirically evaluate it by comparing its performance on simulated data relative to more traditional statistical methods, before applying it to real-world data. Our findings indicate that the application of a suitably flexible machine learning algorithm within DML improves the adjustment for various nonlinear confounding relationships. This advantage enables a departure from traditional functional form assumptions typically necessary in causal effect estimation. However, we demonstrate that the method continues to critically depend on standard assumptions about causal structure and identification. When estimating the effects of air pollution on housing prices in our application, we find that DML estimates are consistently larger than estimates of less flexible methods. From our overall results, we provide actionable recommendations for specific choices researchers must make when applying DML in practice.