2024/07/09 by Arsenii Mustafin, Mustafin, Arsenii, Aleksei Pakharev +5 · 1 citation
Computer Science · #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Matrix Theory and Algorithms #Optimization and Control (math.OC)
paper · pdf · doi:10.48550/arxiv.2407.06712
openalex publication_date 2024/07/09 · openalex created_date 2024/07/12 · openalex updated_date 2026/07/28
The Markov Decision Process (MDP) is a widely used mathematical model for sequential decision-making problems. In this paper, we present a new geometric interpretation of MDPs with a natural normalization procedure that allows us to adjust the value function at each state without altering the advantage of any action with respect to any policy. This procedure enables the development of a novel class of algorithms for solving MDPs that find optimal policies without explicitly computing policy values. The new algorithms we propose for different settings achieve and, in some cases, improve upon state-of-the-art sample complexity results.