2023/01/20 by Sofia Ek, Ek, Sofia, Dave Zachariah +3
Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Statistical Methods and Bayesian Inference #Statistical Methods and Inference #Statistical Methods in Clinical Trials
paper · pdf · doi:10.48550/arxiv.2301.08649
openalex publication_date 2023/01/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01
We consider the problem of evaluating the performance of a decision policy using past observational data. The outcome of a policy is measured in terms of a loss (aka. disutility or negative reward) and the main problem is making valid inferences about its out-of-sample loss when the past data was observed under a different and possibly unknown policy. Using a sample-splitting method, we show that it is possible to draw such inferences with finite-sample coverage guarantees about the entire loss distribution, rather than just its mean. Importantly, the method takes into account model misspecifications of the past policy - including unmeasured confounding. The evaluation method can be used to certify the performance of a policy using observational data under a specified range of credible model assumptions.