vix.ing · top · new · best · stats · spec

Critic Regularized Regression

2020/06/26 by Ziyu Wang, Alexander Novikov, Wang, Ziyu +19 · 12 citations
Computer Science · #Adaptive Dynamic Programming Control #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.2006.15134

openalex publication_date 2020/06/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Offline reinforcement learning (RL), also known as batch RL, offers the prospect of policy optimization from large pre-recorded datasets without online environment interaction. It addresses challenges with regard to the cost of data collection and safety, both of which are particularly pertinent to real-world applications of RL. Unfortunately, most off-policy algorithms perform poorly when learning from a fixed dataset. In this paper, we propose a novel offline RL algorithm to learn policies from data using a form of critic-regularized regression (CRR). We find that CRR performs surprisingly well and scales to tasks with high-dimensional state and action spaces -- outperforming several state-of-the-art offline RL algorithms by a significant margin on a wide range of benchmark tasks.

Citations

Cited by

Related