2018/09/17 by Hyung‐Jin Yoon, Donghwan Lee, Yoon, Hyung-Jin +3
Computer Science · Engineering · #Distributed Sensor Networks and Detection Algorithms #FOS: Computer and information sciences #FOS: Electrical engineering #Fault Detection and Control Systems #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Systems and Control (eess.SY) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1809.06401
openalex publication_date 2018/09/17 · openalex created_date 2022/08/03 · openalex updated_date 2026/07/28
The objective is to study an on-line Hidden Markov model (HMM)\nestimation-based Q-learning algorithm for partially observable Markov decision\nprocess (POMDP) on finite state and action sets. When the full state\nobservation is available, Q-learning finds the optimal action-value function\ngiven the current action (Q function). However, Q-learning can perform poorly\nwhen the full state observation is not available. In this paper, we formulate\nthe POMDP estimation into a HMM estimation problem and propose a recursive\nalgorithm to estimate both the POMDP parameter and Q function concurrently.\nAlso, we show that the POMDP estimation converges to a set of stationary points\nfor the maximum likelihood estimate, and the Q function estimation converges to\na fixed point that satisfies the Bellman optimality equation weighted on the\ninvariant distribution of the state belief determined by the HMM estimation\nprocess.\n