2019/07/12 by Cohen, Samuel N., Treetanthiploet, Tanut
#60G40 #91B32 #91B70 #93E35 #Computational Finance (q-fin.CP) #FOS: Economics and business #FOS: Mathematics #Optimization and Control (math.OC) #Probability (math.PR) #Statistics Theory (math.ST)
paper · doi:10.48550/arxiv.1907.05689
We study dynamic allocation problems for discrete time multi-armed bandits under uncertainty, based on the the theory of nonlinear expectations. We show that, under strong independence of the bandits and with some relaxation in the definition of optimality, a Gittins allocation index gives optimal choices. This involves studying the interaction of our uncertainty with controls which determine the filtration. We also run a simple numerical example which illustrates the interaction between the willingness to explore and uncertainty aversion of the agent when making decisions.