vix.ing · top · new · best · stats

Multi-thresholding Good Arm Identification with Bandit Feedback

2025/03/13 by Jiang, Xuanke, Sherief Hashima, Hashima, Sherief +4
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Stock Market Forecasting Methods

paper · pdf · doi:10.48550/arxiv.2503.10386

openalex publication_date 2025/03/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We consider a good arm identification problem in a stochastic bandit setting with multi-objectives, where each arm i ∈ [K] is associated with a distribution Di defined over RM. For each round t, the player pulls an arm it and receives an M-dimensional reward vector sampled according to Dit. The goal is to find, with high probability, an ε-good arm whose expected reward vector is larger than \bmξ - ε1, where \bmξ is a predefined threshold vector, and the vector comparison is component-wise. We propose the Multi-Thresholding UCB~(MultiTUCB) algorithm with a sample complexity bound. Our bound matches the existing one in the special case where M=1 and ε=0. The proposed algorithm demonstrates superior performance compared to baseline approaches across synthetic and real datasets.

Related