2023/06/21 by Jiacheng Guo, Zihao Li, Guo, Jiacheng +9 · 1 citation
Computer Science · #Bayesian Modeling and Causal Inference #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Algorithms
paper · pdf · doi:10.48550/arxiv.2306.12356
openalex publication_date 2023/06/21 · openalex created_date 2023/06/24 · openalex updated_date 2026/07/28
In this paper, we study representation learning in partially observable Markov Decision Processes (POMDPs), where the agent learns a decoder function that maps a series of high-dimensional raw observations to a compact representation and uses it for more efficient exploration and planning. We focus our attention on the sub-classes of γ-observable and decodable POMDPs, for which it has been shown that statistically tractable learning is possible, but there has not been any computationally efficient algorithm. We first present an algorithm for decodable POMDPs that combines maximum likelihood estimation (MLE) and optimism in the face of uncertainty (OFU) to perform representation learning and achieve efficient sample complexity, while only calling supervised learning computational oracles. We then show how to adapt this algorithm to also work in the broader class of γ-observable POMDPs.