2022/10/11 by O. Bastani, Bastani, O., Y. J. Ma +5 · 2 citations
Decision Sciences · Computer Science · #Advanced Bandit Algorithms Research #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI)
paper · pdf · doi:10.48550/arxiv.2210.05650
In safety-critical applications of reinforcement learning such as healthcare and robotics, it is often desirable to optimize risk-sensitive objectives that account for tail outcomes rather than expected reward. We prove the first regret bounds for reinforcement learning under a general class of risk-sensitive objectives including the popular CVaR objective. Our theory is based on a novel characterization of the CVaR objective as well as a novel optimistic MDP construction.