2025/07/15 by Justin Whitehouse, Whitehouse, Justin, Chen, Qizhao +3 · 3 citations
Economics, Econometrics and Finance · #Climate Change Policy and Economics #Econometrics (econ.EM) #Economic Policies and Impacts #FOS: Computer and information sciences #FOS: Economics and business #FOS: Mathematics #Machine Learning (cs.LG) #Methodology (stat.ME) #Monetary Policy and Economic Impact #Statistics Theory (math.ST)
paper · pdf · doi:10.48550/arxiv.2507.11780
openalex publication_date 2025/07/15 · openalex created_date 2025/10/14 · openalex updated_date 2026/07/31
Constructing confidence intervals for the value of an (unknown) optimal treatment policy is a fundamental problem in causal inference. Insight into the optimal policy value can guide the development of reward-maximizing, individualized treatment regimes. However, because the functional that defines the optimal value is non-differentiable, standard semi-parametric approaches for performing inference fail to be directly applicable. Many existing works circumvent non-differentiability by making the unrealistic assumption of zero probability of treatment non-response, i.e. that every unit responds (either positively or negatively) to an assigned treatment. Further, works that don't circumvent this restriction rely on refitting nuisance models a number of times proportional to the sample size. In this paper, we construct and analyze a simple, softmax smoothing-based estimator for the value of an optimal treatment policy. Our estimator applies in both static and dynamic treatment regimes, only requires fitting a constant number of nuisance models, and is statistically efficient when there is zero probability of non-response to treatment. Also, while our estimator does not require making semi-parametric restrictions, it can exploit them when they exist. We further show how our softmax smoothing approach can be used to estimate general parameters that are specified as a maximum of scores involving nuisance components, and look at conditional Balke and Pearl bounds and L1 calibration error as salient examples.