vix.ing · top · new · best · stats · spec

Thompson Sampling for Gaussian Entropic Risk Bandits

2021/05/14 by Ming Liang Ang, Ang, Ming Liang, Eloise Y. Y. Lim +3
Decision Sciences · Computer Science · #Advanced Bandit Algorithms Research #Reinforcement Learning in Robotics #Machine Learning and Algorithms

paper · pdf · doi:10.48550/arxiv.2105.06960

Abstract

The multi-armed bandit (MAB) problem is a ubiquitous decision-making problem that exemplifies exploration-exploitation tradeoff. Standard formulations exclude risk in decision making. Risknotably complicates the basic reward-maximising objectives, in part because there is no universally agreed definition of it. In this paper, we consider an entropic risk (ER) measure and explore the performance of a Thompson sampling-based algorithm ERTS under this risk measure by providing regret bounds for ERTS and corresponding instance dependent lower bounds.

Related