2025/01/15 by Delattre, Sylvain, Fournier, Nicolas
#FOS: Mathematics #Optimization and Control (math.OC) #Probability (math.PR)
paper · doi:10.48550/arxiv.2501.08800
We consider the Monte-Carlo first visit algorithm, of which the goal is to find the optimal control in a Markov decision process with finite state space and finite number of possible actions. We show its convergence when the discount factor is smaller than 1/2.