2020/08/18 by Umer Siddique, Siddique, Umer, Paul Weng +3 · 4 citations
Computer Science · Social Sciences · Business, Management and Accounting · #Reinforcement Learning in Robotics #Experimental Behavioral Economics Studies #Supply Chain and Inventory Management
paper · pdf · doi:10.48550/arxiv.2008.07773
As the operations of autonomous systems generally affect simultaneously\nseveral users, it is crucial that their designs account for fairness\nconsiderations. In contrast to standard (deep) reinforcement learning (RL), we\ninvestigate the problem of learning a policy that treats its users equitably.\nIn this paper, we formulate this novel RL problem, in which an objective\nfunction, which encodes a notion of fairness that we formally define, is\noptimized. For this problem, we provide a theoretical discussion where we\nexamine the case of discounted rewards and that of average rewards. During this\nanalysis, we notably derive a new result in the standard RL setting, which is\nof independent interest: it states a novel bound on the approximation error\nwith respect to the optimal average reward of that of a policy optimal for the\ndiscounted reward. Since learning with discounted rewards is generally easier,\nthis discussion further justifies finding a fair policy for the average reward\nby learning a fair policy for the discounted reward. Thus, we describe how\nseveral classic deep RL algorithms can be adapted to our fair optimization\nproblem, and we validate our approach with extensive experiments in three\ndifferent domains.\n