vix.ing · top · new · best · stats · spec

Pathwise uniform value in gambling houses and Partially Observable Markov Decision Processes

2015/05/27 by Xavier Venel, Venel, Xavier, Bruno Ziliotto +1
Mathematics · #37A50 #47N10 #90C39 #90C40 #FOS: Mathematics #Optimization and Control (math.OC) #math.OC #msc:37A50 #msc:47N10 #msc:90C39 #msc:90C40

paper · pdf · doi:10.48550/arxiv.1505.07495

arxiv created 2015/09/08 · arxiv updated 2015/09/09

Abstract

In several standard models of dynamic programming (gambling houses, MDPs, POMDPs), we prove the existence of a very robust notion of value for the infinitely repeated problem, namely the pathwise uniform value. This solves two open problems. First, this shows that for any epsilon>0, the decision-maker has a pure strategy sigma which is epsilon-optimal in any n-stage game, provided that n is big enough (this result was only known for behavior strategies, that is, strategies which use randomization). Second, the strategy sigma can be chosen such that under the long-run average payoff criterion (expectation of the liminf of the average payoffs), the decision-maker has more than lim v(n)-epsilon.

Related