vix.ing · top · new · best · stats · spec

Sequential Experimental Design for Transductive Linear Bandits

2019/06/20 by Fiez, Tanner, Jain, Lalit, Jamieson, Kevin +1 · 3 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)

paper · doi:10.48550/arxiv.1906.08399

Abstract

In this paper we introduce the transductive linear bandit problem: given a set of measurement vectors X⊂ ℝd, a set of items Z⊂ ℝd, a fixed confidence δ, and an unknown vector θ∈ ℝd, the goal is to infer argmaxz∈ Z z^\topθ^∗ with probability 1-δ by making as few sequentially chosen noisy measurements of the form x^\topθ as possible. When X=Z, this setting generalizes linear bandits, and when X is the standard basis vectors and Z⊂ \0,1\d, combinatorial bandits. Such a transductive setting naturally arises when the set of measurement vectors is limited due to factors such as availability or cost. As an example, in drug discovery the compounds and dosages X a practitioner may be willing to evaluate in the lab in vitro due to cost or safety reasons may differ vastly from those compounds and dosages Z that can be safely administered to patients in vivo. Alternatively, in recommender systems for books, the set of books X a user is queried about may be restricted to well known best-sellers even though the goal might be to recommend more esoteric titles Z. In this paper, we provide instance-dependent lower bounds for the transductive setting, an algorithm that matches these up to logarithmic factors, and an evaluation. In particular, we provide the first non-asymptotic algorithm for linear bandits that nearly achieves the information theoretic lower bound.

Cited by

Related