vix.ing · top · new · best · stats · spec

How to select predictive models for decision-making or causal inference

2025/01/01 by Matthieu Doutreligne, Gaël Varoquaux · 1 voice · 2 citations
Computer Science · Mathematics · #Advanced Causal Inference Techniques #Bayesian Modeling and Causal Inference #Explainable Artificial Intelligence (XAI)

paper · doi:10.1093/gigascience/giaf016

openalex publication_date 2025/01/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30

Abstract

BACKGROUND: We investigate which procedure selects the most trustworthy predictive model to explain the effect of an intervention and support decision-making. METHODS: We study a large variety of model selection procedures in practical settings: finite samples settings and without a theoretical assumption of well-specified models. Beyond standard cross-validation or internal validation procedures, we also study elaborate causal risks. These build proxies of the causal error using "nuisance" reweighting to compute it on the observed data. We evaluate whether empirically estimated nuisances, which are necessarily noisy, add noise to model selection and compare different metrics for causal model selection in an extensive empirical study based on a simulation and 3 health care datasets based on real covariates. RESULTS: Among all metrics, the mean squared error, classically used to evaluate predictive modes, is worse. Reweighting it with a propensity score does not bring much improvement in most cases. On average, the R-risk, which uses as nuisances a model of mean outcome and propensity scores, leads to the best performances. Nuisance corrections are best estimated with flexible estimators such as a super learner. CONCLUSIONS: When predictive models are used to explain the effect of an intervention, they must be evaluated with different procedures than standard predictive settings, using the R-risk from causal inference.

Citations

Cited by

Discussions

Related