2014/03/17 by Chandrashekar Lakshminarayanan, Shalabh Bhatnagar, Lakshminarayanan, Chandrashekar +1
Computer Science · Medicine · #Reinforcement Learning in Robotics #Adaptive Dynamic Programming Control #Cardiac Valve Diseases and Treatments
paper · pdf · doi:10.48550/arxiv.1403.4179
Markov Decision Processes (MDP) is an useful framework to cast optimal\nsequential decision making problems. Given any MDP the aim is to find the\noptimal action selection mechanism i.e., the optimal policy. Typically, the\noptimal policy (u^*) is obtained by substituting the optimal value-function\n(J^*) in the Bellman equation. Alternately u^* is also obtained by learning\nthe optimal state-action value function Q^* known as the Q value-function.\nHowever, it is difficult to compute the exact values of J^* or Q^* for MDPs\nwith large number of states. Approximate Dynamic Programming (ADP) methods\naddress this difficulty by computing lower dimensional approximations of\nJ^*/Q^*. Most ADP methods employ linear function approximation (LFA), i.e.,\nthe approximate solution lies in a subspace spanned by a family of pre-selected\nbasis functions. The approximation is obtain via a linear least squares\nprojection of higher dimensional quantities and the L2 norm plays an\nimportant role in convergence and error analysis. In this paper, we discuss ADP\nmethods for MDPs based on LFAs in (\min,+) algebra. Here the approximate\nsolution is a (\min,+) linear combination of a set of basis functions whose\nspan constitutes a subsemimodule. Approximation is obtained via a projection\noperator onto the subsemimodule which is different from linear least squares\nprojection used in ADP methods based on conventional LFAs. MDPs are not\n(\min,+) linear systems, nevertheless, we show that the monotonicity property\nof the projection operator helps us to establish the convergence of our ADP\nschemes. We also discuss future directions in ADP methods for MDPs based on the\n(\min,+) LFAs.\n