vix.ing · top · new · best · stats · spec

The Value Iteration Algorithm is Not Strongly Polynomial for Discounted Dynamic Programming

2013/12/19 by Eugene A. Feinberg, Feinberg, Eugene A., Jefferson Huang +1 · 2 citations
Computer Science · Decision Sciences · #Adaptive Dynamic Programming Control #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #FOS: Mathematics #Optimization and Control (math.OC) #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.1312.6832

openalex publication_date 2013/12/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This note provides a simple example demonstrating that, if exact computations are allowed, the number of iterations required for the value iteration algorithm to find an optimal policy for discounted dynamic programming problems may grow arbitrarily quickly with the size of the problem. In particular, the number of iterations can be exponential in the number of actions. Thus, unlike policy iterations, the value iteration algorithm is not strongly polynomial for discounted dynamic programming.

Citations

Cited by

Related