vix.ing · top · new · best · stats · spec

Pure Exploration Bandit Problem with General Reward Functions Depending on Full Distributions

2021/05/08 by Siwei Wang, Wei Chen, Wang, Siwei +1
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #Machine Learning (cs.LG) #Optimization and Search Problems #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.2105.03598

openalex publication_date 2021/05/08 · openalex created_date 2021/05/24 · openalex updated_date 2026/07/28

Abstract

In this paper, we study the pure exploration bandit model on general distribution functions, which means that the reward function of each arm depends on the whole distribution, not only its mean. We adapt the racing framework and LUCB framework to solve this problem, and design algorithms for estimating the value of the reward functions with different types of distributions. Then we show that our estimation methods have correctness guarantee with proper parameters, and obtain sample complexity upper bounds for them. Finally, we discuss about some important applications and their corresponding solutions under our learning framework.

Related