Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
2025/08/05 by Ma, Yi, Tang, Hongyao, Xiao, Chenjun +4 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2508.03194
Abstract
In recent years, the expansion of neural network models and training data has driven remarkable progress in deep learning, particularly in computer vision and natural language processing. This advancement is underpinned by the concept of Scaling Laws, which demonstrates that scaling model parameters and training data enhances learning performance. While these fields have witnessed breakthroughs, such as the development of large language models like GPT-4 and advanced vision models like Midjourney, the application of scaling laws in deep reinforcement learning (DRL) remains relatively unexplored. Despite its potential to improve performance, the integration of scaling laws into DRL for decision making has not been fully realized. This review addresses this gap by systematically analyzing scaling strategies in three dimensions: data, network, and training budget. In data scaling, we explore methods to optimize data efficiency through parallel sampling and data generation, examining the relationship between data volume and learning outcomes. For network scaling, we investigate architectural enhancements, including monolithic expansions, ensemble and MoE methods, and agent number scaling techniques, which collectively enhance model expressivity while posing unique computational challenges. Lastly, in training budget scaling, we evaluate the impact of distributed training, high replay ratios, large batch sizes, and auxiliary training on training efficiency and convergence. By synthesizing these strategies, this review not only highlights their synergistic roles in advancing DRL for decision making but also provides a roadmap for future research. We emphasize the importance of balancing scalability with computational efficiency and outline promising directions for leveraging scaling to unlock the full potential of DRL in various tasks such as robot control, autonomous driving and LLM training.
Citations
- Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
- RePO: Replay-Enhanced Policy Optimization
- Horizon Reduction Makes RL Scalable
- The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks
- Mitigating Plasticity Loss in Continual Reinforcement Learning by Reducing Churn
- Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
- FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
- Qwen3 Technical Report
- Learning to Reason under Off-Policy Guidance
- SmolVLM: Redefining small and efficient multimodal models
- 1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities
- Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers
- Hyperspherical Normalization for Scalable Deep Reinforcement Learning
- Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization
- Value-Based Deep RL Scales Predictably
- Dual Ensembled Multiagent Q-Learning with Hypernet Regularizer
- s1: Simple test-time scaling
- Prioritized Generative Replay
- SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning
- MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL
- Improving Deep Reinforcement Learning by Reducing the Chain Effect of Value and Policy Churn
- SAPG: Split and Aggregate Policy Gradients
- Simplifying Deep Temporal Difference Learning
- Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term Planning
- Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
- Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents
- Higher Replay Ratio Empowers Sample-Efficient Multi-Agent Reinforcement Learning
- HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation
- Dissecting Deep RL with High Update Ratios: Combatting Value Divergence
- Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
- Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning
- Mixtures of Experts Unlock Parameter Scaling for Deep RL
- Offline Actor-Critic Reinforcement Learning Scales to Large Models
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Bridging Evolutionary Algorithms and Reinforcement Learning: A Comprehensive Survey on Hybrid Algorithms
- Keep Various Trajectories: Promoting Exploration of Ensemble Policies in Continuous Control
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- Adaptive Ensemble Q-learning: Minimizing Estimation Bias via Error Feedback
- PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement Learning
- HIPODE: Enhancing Offline Reinforcement Learning with High-Quality Synthetic Data from a Policy-Decoupled Approach
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
- Bigger, Better, Faster: Human-level Atari with human-level efficiency
- Off-Policy RL Algorithms Can be Sample-Efficient for Continuous Control via Sample Multiple Reuse
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- Efficient Deep Reinforcement Learning Requires Regulating Overfitting
- Generative Agents: Interactive Simulacra of Human Behavior
- GPT-4 Technical Report
- Synthetic Experience Replay
- Understanding plasticity in neural networks
- Efficient Online Reinforcement Learning with Offline Data
- Scaling laws for single-agent reinforcement learning
- Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes
- Q-Ensemble for Offline RL: Don't Scale the Ensemble, Scale the Batch Size
- ERL-Re2: Efficient Evolutionary Reinforcement Learning with Shared State Representation and Individual Policy Representation
- Towards A Unified Policy Abstraction Theory and Representation Learning Approach in Markov Decision Processes
- Bootstrapped Transformer for Offline Reinforcement Learning
- Multi-Game Decision Transformers
- MyoSuite -- A contact-rich simulation suite for musculoskeletal motor control
- Towards Applicable Reinforcement Learning: Improving the Generalization and Sample Efficiency with Policy Ensemble
- The Primacy Bias in Deep Reinforcement Learning
- A Generalist Agent
- Training Compute-Optimal Large Language Models
- DNS: Determinantal Point Process Based Neural Network Sampler for Ensemble Reinforcement Learning
- Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning
- Evolutionary Action Selection for Gradient-based Policy Learning
- Aggressive Q-Learning with Ensembles: Achieving Both High Sample Efficiency and High Asymptotic Performance
- URLB: Unsupervised Reinforcement Learning Benchmark
- Dropout Q-Functions for Doubly Efficient Reinforcement Learning
- Large Batch Experience Replay
- Uncertainty-Based Offline Reinforcement Learning with Diversified Q-Ensemble
- HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action Representation
- Learning to Prompt for Vision-Language Models
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning
- Offline-to-Online Reinforcement Learning via Balanced Replay and Pessimistic Q-Ensemble
- PlayVirtual: Augmenting Cycle-Consistent Virtual Trajectories for Reinforcement Learning
- Efficient Continuous Control with Double Actors and Regularized Critics
- Ensemble Bootstrapping for Q-Learning
- Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
- Cooperative Heterogeneous Deep Reinforcement Learning
- Scaling Laws for Autoregressive Generative Modeling
- Masked Contrastive Representation Learning for Reinforcement Learning
- Phasic Policy Gradient
- Sample-Efficient Automated Deep Reinforcement Learning
- Data-Efficient Reinforcement Learning with Self-Predictive\n Representations
- Revisiting Fundamentals of Experience Replay
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep\n Reinforcement Learning
- Learning Invariant Representations for Reinforcement Learning without Reconstruction
- Experience Augmentation: Boosting and Accelerating Off-Policy Multi-Agent Reinforcement Learning
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile Critics
- Reinforcement Learning with Augmented Data
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels
- D4RL: Datasets for Deep Data-Driven Reinforcement Learning
- CURL: Contrastive Unsupervised Representations for Reinforcement Learning
- Can Increasing Input Dimensionality Improve Deep Reinforcement Learning?
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learning
- Scaling Laws for Neural Language Models
- Dota 2 with Large Scale Deep Reinforcement Learning
- Momentum Contrast for Unsupervised Visual Representation Learning
- Momentum Contrast for Unsupervised Visual Representation Learning
- Stabilizing Transformers for Reinforcement Learning
- Improving Sample Efficiency in Model-Free Reinforcement Learning from Images
- Unsupervised State Representation Learning in Atari
- Q-Learning for Continuous Actions with Cross-Entropy Guided Policies
- CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity
- Learning Action Representations for Reinforcement Learning
- ACE: An Actor Ensemble Algorithm for Continuous Control with Tree Search
- Representation Learning with Contrastive Predictive Coding
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
- Evolution-Guided Policy Gradient in Reinforcement Learning
- Accelerated Methods for Deep Reinforcement Learning
- Addressing Function Approximation Error in Actor-Critic Methods
- Distributed Prioritized Experience Replay
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- DeepMind Control Suite
- Rainbow: Combining Improvements in Deep Reinforcement Learning
- Proximal Policy Optimization Algorithms
- Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
- Online Meta-learning by Parallel Algorithm Competition
- Reinforcement Learning with Unsupervised Auxiliary Tasks
- Deep Exploration via Bootstrapped DQN
- Asynchronous Methods for Deep Reinforcement Learning
- Deep Reinforcement Learning with Double Q-learning
- Continuous control with deep reinforcement learning
- Playing Atari with Deep Reinforcement Learning
- Path Integral Policy Improvement with Covariance Matrix Adaptation
- Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
- Reinforcement Learning: An Introduction
Cited by
Related