vix.ing · top · new · best · stats · spec

Provably Safe Reinforcement Learning with Step-wise Violation Constraints

2023/02/13 by Nuoya Xiong, Xiong, Nuoya, Yihan du +3 · 2 citations
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.2302.06064

openalex publication_date 2023/02/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In this paper, we investigate a novel safe reinforcement learning problem with step-wise violation constraints. Our problem differs from existing works in that we consider stricter step-wise violation constraints and do not assume the existence of safe actions, making our formulation more suitable for safety-critical applications which need to ensure safety in all decision steps and may not always possess safe actions, e.g., robot control and autonomous driving. We propose a novel algorithm SUCBVI, which guarantees \widetildeO(√(ST)) step-wise violation and \widetildeO(√(H3SAT)) regret. Lower bounds are provided to validate the optimality in both violation and regret performance with respect to S and T. Moreover, we further study a novel safe reward-free exploration problem with step-wise violation constraints. For this problem, we design an (ε,δ)-PAC algorithm SRF-UCRL, which achieves nearly state-of-the-art sample complexity \widetildeO(((S2AH2)/(ε)+(H4SA)/(ε2))(log(\frac1δ)+S)), and guarantees \widetildeO(√(ST)) violation during the exploration. The experimental results demonstrate the superiority of our algorithms in safety performance, and corroborate our theoretical results.

Cited by

Related