2025/02/19 by Mifrani, Anas, Noll, Dominikus
#90C05 #90C29 #90C40 #FOS: Mathematics #Optimization and Control (math.OC)
paper · doi:10.48550/arxiv.2502.13697
We propose a vector linear programming formulation for a non-stationary, finite-horizon Markov decision process with vector-valued rewards. Pareto efficient policies are shown to correspond to efficient solutions of the linear program, and vector linear programming theory allows us to fully characterize deterministic efficient policies. An algorithm for enumerating all efficient deterministic policies is presented then tested numerically in an engineering application.