vix.ing · top · new · best · stats · spec

Contextual Procurement Auctions with Bandit Learning

2026/07/27 by Yiling Chen, Shi Feng, Sadie Zhao
#cs.GT #cs.LG

paper · pdf

Abstract

We study repeated procurement auctions in which producers have private costs and the platform must learn the context-dependent value of selecting each producer. We evaluate performance by welfare regret: the cumulative loss in total surplus relative to the full-information efficient rule that knows the context-dependent values and true producer costs. The natural UCB allocation rule achieves \widetilde O(√(ngT)) welfare regret under truthful bids, but its adaptive, bid-dependent learning path does not by itself ensure truthfulness. To obtain exact incentives, we first design a bid-independent explore-then-commit mechanism with empirical threshold payments; it is dominant-strategy truthful and has \widetilde O((ng)1/3T2/3) regret. We then introduce frozen-payment UCB, which estimates payments from initial bid-independent exploration but continues allocation learning by UCB. Under a truthful-path margin condition, the frozen-payment UCB is approximately truthful with an average per-round deviation gain \widetilde O(T-1/4) for fixed n, g. Under truthful bidding, it achieves \widetilde O(√(ngT)) welfare regret, matching the UCB rate. A lower bound shows that this regret-incentive tradeoff is tight within the frozen critical-payment class.

Citations

Related