vix.ing · top · new · best · stats · spec

On Value Iteration Convergence in Connected MDPs

2024/06/13 by Arsenii Mustafin, Alex Olshevsky, Mustafin, Arsenii +3
Computer Science · Engineering · #Advanced Control Systems Optimization #FOS: Computer and information sciences #Machine Learning (cs.LG) #Optimization and Variational Analysis

paper · pdf · doi:10.48550/arxiv.2406.09592

openalex publication_date 2024/06/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This paper establishes that an MDP with a unique optimal policy and ergodic associated transition matrix ensures the convergence of various versions of the Value Iteration algorithm at a geometric rate that exceeds the discount factor γ for both discounted and average-reward criteria.

Related