vix.ing · top · new · best · stats · spec

Kernel weighted importance sampling for off-policy evaluation in contextual bandits

2026/07/16 by Joshua Spear, Matthieu Komorowski, Rebecca Pope +2
#cs.LG

paper · pdf

Abstract

This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically outperform strong baselines (including vanilla weighted importance sampling), particularly under complex conditions including behaviour policy miss-specification. The benefit of Kernel-WIS is derived from combining the bounded property of vanilla weighted importance sampling with the linearity of vanilla importance sampling.

Related