2015/07/04 by Mahmoud El Chamie, Chamie, Mahmoud El, Behçet Açıkmeşe +1
Computer Science · Decision Sciences · #90C40 #Advanced Bandit Algorithms Research #FOS: Electrical engineering #FOS: Mathematics #G.1.6 #G.3 #I.2.8 #Optimization and Control (math.OC) #Optimization and Search Problems #Reinforcement Learning in Robotics #Simulation Techniques and Applications #Software Reliability and Analysis Research #Systems and Control (eess.SY) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1507.01151
openalex publication_date 2015/07/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Markov Decision Processes (MDPs) have been used to formulate many decision-making problems in science and engineering. The objective is to synthesize the best decision (action selection) policies to maximize expected rewards (or minimize costs) in a given stochastic dynamical environment. In this paper, we extend this model by incorporating additional information that the transitions due to actions can be sequentially observed. The proposed model benefits from this information and produces policies with better performance than those of standard MDPs. The paper also presents an efficient offline linear programming based algorithm to synthesize optimal policies for the extended model.