vix.ing · top · new · best · stats · spec

Sufficiency of Markov Policies for Continuous-Time Jump Markov Decision\n Processes

2020/03/03 by Eugene A. Feinberg, Feinberg, Eugene A., Manasa Mandava +3
Decision Sciences · #60J25 #FOS: Mathematics #Optimization and Control (math.OC) #Primary: 90C40 #Simulation Techniques and Applications #secondary: 90C39

paper · pdf · doi:10.48550/arxiv.2003.01342

openalex publication_date 2020/03/03 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28

Abstract

This paper extends to Continuous-Time Jump Markov Decision Processes (CTJMDP)\nthe classic result for Markov Decision Processes stating that, for a given\ninitial state distribution, for every policy there is a (randomized) Markov\npolicy, which can be defined in a natural way, such that at each time instance\nthe marginal distributions of state-action pairs for these two policies\ncoincide. It is shown in this paper that this equality takes place for a CTJMDP\nif the corresponding Markov policy defines a nonexplosive jump Markov process.\nIf this Markov process is explosive, then at each time instance the marginal\nprobability, that a state-action pair belongs to a measurable set of\nstate-action pairs, is not greater for the described Markov policy than the\nsame probability for the original policy. These results are used in this paper\nto prove that for expected discounted total costs and for average costs per\nunit time, for a given initial state distribution, for each policy for a CTJMDP\nthe described a Markov policy has the same or better performance.\n

Related