vix.ing · top · new · best · stats

Efficient Inference in Markov Control Problems

2012/02/14 by Thomas Furmston, David Barber, Furmston, Thomas +1 · 2 citations
Computer Science · Decision Sciences · Mathematics · #Advanced Bandit Algorithms Research #Algorithm #Approximate inference #Artificial intelligence #Computer science #Extension (predicate logic) #Horizon #Inference #Machine learning #Markov chain #Markov decision process #Markov model #Markov process #Mathematical optimization #Mathematics #Optimization and Search Problems #Reinforcement Learning in Robotics #Trajectory

paper · pdf · doi:10.48550/arxiv.1202.3720

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2012/02/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Markov control algorithms that perform smooth, non-greedy updates of the policy have been shown to be very general and versatile, with policy gradient and Expectation Maximisation algorithms being particularly popular. For these algorithms, marginal inference of the reward weighted trajectory distribution is required to perform policy updates. We discuss a new exact inference algorithm for these marginals in the finite horizon case that is more efficient than the standard approach based on classical forward-backward recursions. We also provide a principled extension to infinite horizon Markov Decision Problems that explicitly accounts for an infinite horizon. This extension provides a novel algorithm for both policy gradients and Expectation Maximisation in infinite horizon problems.

Citations

Cited by

Related