2015/05/04 by Denis Nekipelov, Vasilis Syrgkanis, Nekipelov, Denis +4 · 3 citations
Business, Management and Accounting · Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Artificial intelligence #Auction Theory and Applications #Best response #Computer Science and Game Theory (cs.GT) #Computer science #Consumer Market Behavior and Pricing #Economics #Epsilon-equilibrium #Equilibrium selection #FOS: Computer and information sciences #Game theory #Inference #Machine learning #Mathematical economics #Nash equilibrium #Regret #Repeated game #cs.GT
paper · pdf · doi:10.48550/arxiv.1505.00720
arxiv created 2015/05/04 · openalex publication_date 2015/05/04 · arxiv updated 2015/05/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
The main goal of this paper is to develop a theory of inference of player valuations from observed data in the generalized second price auction without relying on the Nash equilibrium assumption. Existing work in Economics on inferring agent values from data relies on the assumption that all participant strategies are best responses of the observed play of other players, i.e. they constitute a Nash equilibrium. In this paper, we show how to perform inference relying on a weaker assumption instead: assuming that players are using some form of no-regret learning. Learning outcomes emerged in recent years as an attractive alternative to Nash equilibrium in analyzing game outcomes, modeling players who haven't reached a stable equilibrium, but rather use algorithmic learning, aiming to learn the best way to play from previous observations. In this paper we show how to infer values of players who use algorithmic learning strategies. Such inference is an important first step before we move to testing any learning theoretic behavioral model on auction data. We apply our techniques to a dataset from Microsoft's sponsored search ad auction system.