vix.ing · top · new · best · stats · spec

Rethinking the Discount Factor in Reinforcement Learning: A Decision\n Theoretic Approach

2019/02/07 by Silviu Pitis, Pitis, Silviu · 2 citations
Decision Sciences · Economics, Econometrics and Finance · #Artificial Intelligence (cs.AI) #Decision-Making and Behavioral Economics #Economic and Environmental Valuation #FOS: Computer and information sciences #Machine Learning (cs.LG)

paper · pdf · doi:10.48550/arxiv.1902.02893

openalex publication_date 2019/02/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Reinforcement learning (RL) agents have traditionally been tasked with\nmaximizing the value function of a Markov decision process (MDP), either in\ncontinuous settings, with fixed discount factor \γ < 1, or in episodic\nsettings, with \γ = 1. While this has proven effective for specific tasks\nwith well-defined objectives (e.g., games), it has never been established that\nfixed discounting is suitable for general purpose use (e.g., as a model of\nhuman preferences). This paper characterizes rationality in sequential decision\nmaking using a set of seven axioms and arrives at a form of discounting that\ngeneralizes traditional fixed discounting. In particular, our framework admits\na state-action dependent "discount" factor that is not constrained to be less\nthan 1, so long as there is eventual long run discounting. Although this\nbroadens the range of possible preference structures in continuous settings, we\nshow that there exists a unique "optimizing MDP" with fixed \γ < 1 whose\noptimal value function matches the true utility of the optimal policy, and we\nquantify the difference between value and utility for suboptimal policies. Our\nwork can be seen as providing a normative justification for (a slight\ngeneralization of) Martha White's RL task formalism (2017) and other recent\ndepartures from the traditional RL, and is relevant to task specification in\nRL, inverse RL and preference-based RL.\n

Cited by

Related