vix.ing · top · new · best · stats

Verification of Markov Decision Processes using Learning Algorithms

2014/02/10 by Tomǎš Brázdil, Tomáš Brázdil, Krishnendu Chatterjee +14 · 10 citations
Computer Science · Mathematics · #Algorithm #Artificial intelligence #Bayesian Modeling and Causal Inference #Bounded function #Computer science #Extension (predicate logic) #FOS: Computer and information sciences #Formal Methods in Verification #Heuristic #Logic in Computer Science (cs.LO) #Machine Learning and Algorithms #Machine learning #Markov chain #Markov decision process #Markov process #Mathematical optimization #Mathematics #Model checking #Probabilistic logic #Reachability #State space #Upper and lower bounds #cs.LO

paper · pdf · doi:10.48550/arxiv.1402.2967

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2014/02/10 · arxiv created 2015/03/30 · arxiv updated 2015/03/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We present a general framework for applying machine-learning algorithms to the verification of Markov decision processes (MDPs). The primary goal of these techniques is to improve performance by avoiding an exhaustive exploration of the state space. Our framework focuses on probabilistic reachability, which is a core property for verification, and is illustrated through two distinct instantiations. The first assumes that full knowledge of the MDP is available, and performs a heuristic-driven partial exploration of the model, yielding precise lower and upper bounds on the required probability. The second tackles the case where we may only sample the MDP, and yields probabilistic guarantees, again in terms of both the lower and upper bounds, which provides efficient stopping criteria for the approximation. The latter is the first extension of statistical model-checking for unbounded properties in MDPs. In contrast with other related approaches, we do not restrict our attention to time-bounded (finite-horizon) or discounted properties, nor assume any particular properties of the MDP. We also show how our techniques extend to LTL objectives. We present experimental results showing the performance of our framework on several examples.

Cited by

Related