Bigger, Better, Faster: Human-level Atari with human-level efficiency
2023/05/30 by Max Schwarzer, Schwarzer, Max, Johan Obando-Ceron +10 · 1 voice · 52 citations
Computer Science · Mathematics · #Anomaly Detection Techniques and Applications #Artificial intelligence #Artificial neural network #Benchmark (surveying) #Code (set theory) #Computer science #Deep neural networks #Human Pose and Action Recognition #Machine learning #Mathematics #Reinforcement Learning in Robotics #Sample (material) #Scaling #Tree (set theory) #Value (mathematics) #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2305.19452
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/05/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We introduce a value-based RL agent, which we call BBF, that achieves super-human performance in the Atari 100K benchmark. BBF relies on scaling the neural networks used for value estimation, as well as a number of other design choices that enable this scaling in a sample-efficient manner. We conduct extensive analyses of these design choices and provide insights for future work. We end with a discussion about updating the goalposts for sample-efficient RL research on the ALE. We make our code and data publicly available at https://github.com/google-research/google-research/tree/master/biggerbetterfaster.
Cited by
- Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
- Relative Entropy Pathwise Policy Optimization
- Mind the GAP! The Challenges of Scale in Pixel-based Deep Reinforcement Learning
- Value-Based Deep RL Scales Predictably
- Recurrent Off-Policy Deep Reinforcement Learning Doesn't Have to be Slow
- Learning Sim-to-Real Humanoid Locomotion in 15 Minutes
- Staggered Environment Resets Improve Massively Parallel On-Policy Reinforcement Learning
- Representation Learning Enables Scalable Multitask Deep Reinforcement Learning
- Natural Building Blocks for Structured World Models: Theory, Evidence, and Scaling
- D2 Actor Critic: Diffusion Actor Meets Distributional Critic
- Human-Like Goalkeeping in a Realistic Football Simulation: a Sample-Efficient Reinforcement Learning Approach
- The Formalism-Implementation Gap in Reinforcement Learning Research
- Guardian: Decoupling Exploration from Safety in Reinforcement Learning
- Toward Agents That Reason About Their Computation
- Confounding Robust Deep Reinforcement Learning: A Causal Approach
- Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
- Rethinking the Role of Dynamic Sparse Training for Scalable Deep Reinforcement Learning
- Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
- DyMoDreamer: World Modeling with Dynamic Modulation
- ZeroSiam: An Efficient Siamese for Test-Time Entropy Optimization without Collapse
- Ratatouille: Imitation Learning Ingredients for Real-world Social Robot Navigation
- On the Limits of Tabular Hardness Metrics for Deep RL: A Study with the Pharos Benchmark
- Compute-Optimal Scaling for Value-Based Deep RL
- Sample-efficient LLM Optimization with Reset Replay
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- Is Exploration or Optimization the Problem for Deep Reinforcement Learning?
- What Can Grokking Teach Us About Learning Under Nonstationarity?
- Balancing Expressivity and Robustness: Constrained Rational Activations for Reinforcement Learning
- Online Training and Pruning of Deep Reinforcement Learning Networks
- EXPO: Stable Reinforcement Learning with Expressive Policies
- Accurate and Efficient World Modeling with Masked Latent Transformers
- A Forget-and-Grow Strategy for Deep Reinforcement Learning Scaling in Continuous Control
- Reliability-Adjusted Prioritized Experience Replay
- Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning
- Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning
- The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning
- Growing with Experience: Growing Neural Networks in Deep Reinforcement Learning
- Bridging the Performance Gap Between Target-Free and Target-Based Reinforcement Learning
- The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks
- Simple, Good, Fast: Self-Supervised World Models Free of Baggage
- Mitigating Plasticity Loss in Continual Reinforcement Learning by Reducing Churn
- AXIOM: Learning to Play Games in Minutes with Expanding Object-Centric Models
- Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
- Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning
- FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
- Hadamax Encoding: Elevating Performance in Model-Free Atari
- Stable Deep Reinforcement Learning via Isotropic Gaussian Representations
- Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities
- Revisiting TD Target Aggregation under Uncertainty in Q-Learning
- Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control
- Zero-Shot Whole-Body Humanoid Control via Behavioral Foundation Models
- Plasticity-Aware Mixture of Experts for Learning Under QoE Shifts in Adaptive Video Streaming
Discussions
Related