2004/08/11 by Frédéric Dambreville, Dambreville, Frederic
Computer Science · #Algorithms and Data Compression #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #FOS: Mathematics #General Mathematics (math.GM) #Machine Learning (cs.LG) #Machine Learning and Algorithms #Reinforcement Learning in Robotics
paper · doi:10.48550/arxiv.math/0408146
openalex publication_date 2004/08/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper, we are interested in optimal decisions in a partially observable Markov universe. Our viewpoint departs from the dynamic programming viewpoint: we are directly approximating an optimal strategic tree depending on the observation. This approximation is made by means of a parameterized probabilistic law. In this paper, a particular family of hidden Markov models, with input and output, is considered as a learning framework. A method for optimizing the parameters of these HMMs is proposed and applied. This optimization method is based on the cross-entropic principle.