vix.ing · top · new · best · stats

Kernelized Offline Contextual Dueling Bandits

2023/07/21 by Viraj Mehta, Mehta, Viraj, Ojash Neopane +9 · 1 citation
Computer Science · Decision Sciences · Mathematics · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #Artificial intelligence #Computer science #Economics #Empirical evidence #FOS: Computer and information sciences #Function (biology) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Machine learning #Mathematics #Microeconomics #Order (exchange) #Preference #Regret #Reinforcement Learning in Robotics #Reinforcement learning #Style (visual arts) #Upper and lower bounds

paper · pdf · doi:10.48550/arxiv.2307.11288

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2023/07/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Preference-based feedback is important for many applications where direct evaluation of a reward function is not feasible. A notable recent example arises in reinforcement learning from human feedback on large language models. For many of these applications, the cost of acquiring the human feedback can be substantial or even prohibitive. In this work, we take advantage of the fact that often the agent can choose contexts at which to obtain human feedback in order to most efficiently identify a good policy, and introduce the offline contextual dueling bandit setting. We give an upper-confidence-bound style algorithm for this setting and prove a regret bound. We also give empirical confirmation that this method outperforms a similar strategy that uses uniformly sampled contexts.

Cited by

Related