A Brief Survey of Deep Reinforcement Learning
2017/08/19 by Kai Arulkumaran, Marc Peter Deisenroth, Miles Brundage +2 · 1 voice · 4,362 citations
Computer Science · Mathematics · #Artificial intelligence #Artificial neural network #Asynchronous communication #Computer science #Deep learning #Field (mathematics) #Machine learning #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Reinforcement learning #Robot #Robotics #Telecommunications #Visual Attention and Saliency Detection #cs.AI #cs.CV #cs.LG #stat.ML
paper · pdf · doi:10.1109/msp.2017.2743240
published in IEEE Signal Processing Magazine 34(6), 26-38 (Institute of Electrical and Electronics Engineers) · IEEE Signal Processing Magazine, Special Issue on Deep Learning for Image Understanding (arXiv extended version)
arxiv created 2017/09/28 · openalex publication_date 2017/11/01 · arxiv updated 2017/11/15 · openalex created_date 2020/11/23 · openalex updated_date 2026/08/05
Abstract
Deep reinforcement learning is poised to revolutionise the field of AI and represents a step towards building autonomous systems with a higher level understanding of the visual world. Currently, deep learning is enabling reinforcement learning to scale to problems that were previously intractable, such as learning to play video games directly from pixels. Deep reinforcement learning algorithms are also applied to robotics, allowing control policies for robots to be learned directly from camera inputs in the real world. In this survey, we begin with an introduction to the general field of reinforcement learning, then progress to the main streams of value-based and policy-based methods. Our survey will cover central algorithms in deep reinforcement learning, including the deep Q-network, trust region policy optimisation, and asynchronous advantage actor-critic. In parallel, we highlight the unique advantages of deep neural networks, focusing on visual understanding via reinforcement learning. To conclude, we describe several current areas of research within the field.
Cited by
- Artificial intelligence for quantum computing
- Deep neural network models of emotion understanding
- Reinforcement learning for control design of uncertain polytopic systems
- Memory Instance Gated Transformer Reinforcement Learning for Portfolio Management
- Strong mixed-integer programming formulations for trained neural networks
- Wireless Networks Design in the Era of Deep Learning: Model-Based, AI-Based, or Both?
- Algebras of actions in an agent's representations of the world
- Decentralized Deep Reinforcement Learning for Network Level Traffic Signal Control
- Reinforcing Medical Image Classifier to Improve Generalization on Small Datasets
- Large-scale agent-based modelling of street robbery using graphical processing units and reinforcement learning
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
- Optimized Conflict Management for Urban Air Mobility Using Swarm UAV Networks
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- LeMat-GenBench: A Unified Evaluation Framework for Crystal Generative Models
- Digital Twin-based Control Co-Design of Full Vehicle Active Suspensions via Deep Reinforcement Learning
- Empowering Things with Intelligence: A Survey of the Progress, Challenges, and Opportunities in Artificial Intelligence of Things
- Emergent Coordination and Phase Structure in Independent Multi-Agent Reinforcement Learning
- DRLViz: Understanding Decisions and Memory in Deep Reinforcement Learning
- Federated Learning in Mobile Edge Networks: A Comprehensive Survey
- Assessing Generalization in Deep Reinforcement Learning
- A Survey of Deep Reinforcement Learning in Video Games
- A Survey of Multi-Access Edge Computing in 5G and Beyond: Fundamentals, Technology Integration, and State-of-the-Art
- Towards Obstacle-Avoiding Control of Planar Snake Robots Exploring Neuro-Evolution of Augmenting Topologies
- Deep reinforcement learning-based spacecraft attitude control with pointing keep-out constraint
- Graph Neural Networks, Deep Reinforcement Learning and Probabilistic Topic Modeling for Strategic Multiagent Settings
- Edge Artificial Intelligence for 6G: Vision, Enabling Technologies, and Applications
- Intelligent Optimization of Multi-Parameter Micromixers Using a Scientific Machine Learning Framework
- Deep Reinforcement Learning for Dynamic Origin-Destination Matrix Estimation in Microscopic Traffic Simulations Considering Credit Assignment
- Value of Information-Enhanced Exploration in Bootstrapped DQN
- Reinforcement learning based data assimilation for unknown state model
- When Semantics Connect the Swarm: LLM-Driven Fuzzy Control for Cooperative Multi-Robot Underwater Coverage
- Single-agent Reinforcement Learning Model for Regional Adaptive Traffic Signal Control
- Robust Single-Agent Reinforcement Learning for Regional Traffic Signal Control Under Demand Fluctuations
- Federated Learning in Mobile Edge Networks: A Comprehensive Survey
- Artificial Intelligence for Safety-Critical Systems in Industrial and Transportation Domains: A Survey
- Deep Reinforcement Learning for Autonomous Internet of Things: Model, Applications and Challenges
- Applications of machine learning in gravitational-wave research with current interferometric detectors
- Edge Artificial Intelligence for 6G: Vision, Enabling Technologies, and Applications
- Dynamics of striatal action selection and reinforcement learning
- Model-Free Design of Stochastic LQR Controller from Reinforcement Learning and Primal-Dual Optimization Perspective
- An Invitation to Deep Reinforcement Learning
- Hybrid Modeling, Sim-to-Real Reinforcement Learning, and Large Language Model Driven Control for Digital Twins
- Guardian: Decoupling Exploration from Safety in Reinforcement Learning
- FlowCritic: Bridging Value Estimation with Flow Matching in Reinforcement Learning
- Video-to-Video Synthesis
- A New Approach for Tactical Decision Making in Lane Changing: Sample Efficient Deep Q Learning with a Safety Feedback Reward
- Deep Learning based Wireless Resource Allocation with Application to Vehicular Networks
- Improved Robustness of Deep Reinforcement Learning for Control of Time-Varying Systems by Bounded Extremum Seeking
- Transformer Based Reinforcement Learning For Games
- A Comprehensive Survey of Incentive Mechanism for Federated Learning
- Bayesian Optimization Meets Riemannian Manifolds in Robot Learning
- Reinforcement Learning for Traversing Chemical Structure Space: Optimizing Transition States and Minimum Energy Paths of Molecules
- Energy-Efficient Design for a NOMA assisted STAR-RIS Network with Deep Reinforcement Learning
- Prioritizing Latency with Profit: A DRL-Based Admission Control for 5G Network Slices
- Safety behavior abstraction and model evolution in autonomous driving
- Interpretable deep learning: interpretation, interpretability, trustworthiness, and beyond
- A Comprehensive Survey: Evaluating the Efficiency of Artificial Intelligence and Machine Learning Techniques on Cyber Security Solutions
- Fine-Tuning LLMs to Analyze Multiple Dimensions of Code Review: A Maximum Entropy Regulated Long Chain-of-Thought Approach
- Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data Augmentation
- SavviDriver: model-based framework for game-based testing of autonomous vehicles in diverse multi-agent traffic scenarios
- Structurally Flexible Neural Networks: Evolving the Building Blocks for General Agents
- Artificial intelligence in the forest products supply chain: current applications and open challenges
- Reinforcement Learning-Based Heuristics to Guide Domain-Independent Dynamic Programming
- A Reinforcement Learning Approach for an IRS-assisted NOMA Network
- AI Agent Access (A\3) Network: An Embodied, Communication-Aware Multi-Agent Framework for 6G Coverage
- SeaPearl: A Constraint Programming Solver guided by Reinforcement Learning
- HypeMARL: Multi-Agent Reinforcement Learning For High-Dimensional, Parametric, and Distributed Systems
- Modeling Worlds in Text
- Multi-Vehicle Routing Problems with Soft Time Windows: A Multi-Agent Reinforcement Learning Approach
- Monte Carlo tree search with spectral expansion for planning with dynamical systems
- Learning Knowledge Graph-based World Models of Textual Environments
- Deep Learning-based Techniques for Integrated Sensing and Communication Systems: State-of-the-Art, Challenges, and Opportunities
- BuildingGym: An open-source toolbox for AI-based building energy management using reinforcement learning
- Deep Reinforcement Learning-Assisted Component Auto-Configuration of Differential Evolution Algorithm for Constrained Optimization: A Foundation Model
- Deep Learning for Insider Threat Detection: Review, Challenges and Opportunities
- Using AI to Optimize Patient Transfer and Resource Utilization During Mass-Casualty Incidents: A Simulation Platform
- The NetHack Learning Environment
- Learning a generative model for robot control using visual feedback
- Communication-Efficient Edge AI: Algorithms and Systems
- DRLViz: Understanding Decisions and Memory in Deep Reinforcement Learning
- HyperNCA: Growing Developmental Networks with Neural Cellular Automata
- A Dynamical Systems Framework for Reinforcement Learning Safety and Robustness Verification
- Optimal control of batch processes via a deterministic Q-learning method
- Review of deep learning: concepts, CNN architectures, challenges, applications, future directions
- Scientific Machine Learning Through Physics–Informed Neural Networks: Where we are and What’s Next
- Fuel Consumption in Platoons: A Literature Review
- A Review On Safe Reinforcement Learning Using Lyapunov and Barrier Functions
- Generalized Operating Procedure for Deep Learning: an Unconstrained Optimal Design Perspective
- Cellular UAV-to-Device Communications: Trajectory Design and Mode Selection by Multi-agent Deep Reinforcement Learning
- Conservative data-driven finite element framework with adaptive hp-refinement for diffusion problems with material uncertainty
- Auto-Agent-Distiller: Towards Efficient Deep Reinforcement Learning Agents via Neural Architecture Search
- A3C-S: Automated Agent Accelerator Co-Search towards Efficient Deep Reinforcement Learning
- Language Model Guided Reinforcement Learning in Quantitative Trading
- Task-Relevant Object Discovery and Categorization for Playing First-person Shooter Games
- Reconstructing 4D Spatial Intelligence: A Survey
- RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow
- Model-Aided Wireless Artificial Intelligence: Embedding Expert Knowledge in Deep Neural Networks Towards Wireless Systems Optimization
- Deep Representation Learning in Speech Processing: Challenges, Recent Advances, and Future Trends
- Episodic Self-Imitation Learning with Hindsight
- Deep reinforcement learning for the real-time inventory rack storage assignment and replenishment problem
- Learning to Communicate in Multi-Agent Reinforcement Learning for Autonomous Cyber Defence
- Federated Reinforcement Learning in Heterogeneous Environments
- Predictive maintenance enabled by machine learning: Use cases and challenges in the automotive industry
- On Deep Learning for Radio Resource Management in A Non-stationary Radio Environment
- A Survey of Deep Reinforcement Learning in Recommender Systems: A Systematic Review and Future Directions
- A Survey on Human-aware Robot Navigation
- Optimal Operating Strategy for PV-BESS Households: Balancing Self-Consumption and Self-Sufficiency
- How is AI Helping Humans Understand the Vocal Communication of Other Species: Recent Advancements in the Use of AI for Acoustic Monitoring of Wildlife
- Physics-Informed Neural Networks For Semiconductor Film Deposition: A Review
- Adaptive Policy Synchronization for Scalable Reinforcement Learning
- Finetuning Deep Reinforcement Learning Policies with Evolutionary Strategies for Control of Underactuated Robots
- A review of mobile robot motion planning methods: from classical motion planning workflows to reinforcement learning-based architectures
- Deep Actor-Critic Learning for Distributed Power Control in Wireless Mobile Networks
- OpenFed: A Comprehensive and Versatile Open-Source Federated Learning Framework
- Intelligent Control of Spacecraft Reaction Wheel Attitude Using Deep Reinforcement Learning
- Deep Reinforcement Learning in Applied Control: Challenges, Analysis, and Insights
- Learning to Expand: Reinforced Pseudo-relevance Feedback Selection for Information-seeking Conversations
- DiffNMR: Diffusion Models for Nuclear Magnetic Resonance Spectra Elucidation
- Deep Learning Techniques for Future Intelligent Cross-Media Retrieval
- Constrained Policy Gradient Method for Safe and Fast Reinforcement Learning: a Neural Tangent Kernel Based Approach
- Generating Novelty in Open-World Multi-Agent Strategic Board Games
- Active Screening for Recurrent Diseases: A Reinforcement Learning Approach
- Indoor thermal comfort management: A Bayesian machine-learning approach to data denoising and dynamics prediction of HVAC systems
- On Hyper-parameter Tuning for Stochastic Optimization Algorithms
- Sample-Efficient Reinforcement Learning with Maximum Entropy Mellowmax Episodic Control
- Zeroth-order Deterministic Policy Gradient
- Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning
- MARL-MambaContour: Unleashing Multi-Agent Deep Reinforcement Learning for Active Contour Optimization in Medical Image Segmentation
- Conservative data-driven finite element framework
- Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities
- Avoiding Catastrophe: Active Dendrites Enable Multi-Task Learning in Dynamic Environments
- AI-based Approach in Early Warning Systems: Focus on Emergency Communication Ecosystem and Citizen Participation in Nordic Countries
- Energy-Based Transfer for Reinforcement Learning
- Hybrid Near-Far Field 6D Movable Antenna Design Exploiting Directional Sparsity and Deep Learning
- PNCS:Power-Norm Cosine Similarity for Diverse Client Selection in Federated Learning
- MARLFuzz: industrial control protocols fuzzing based on multi-agent reinforcement learning
- Learning Swing-up Maneuvers for a Suspended Aerial Manipulation Platform in a Hierarchical Control Framework
- Social physics
- CARoL: Context-aware Adaptation for Robot Learning
- ReLMoGen: Leveraging Motion Generation in Reinforcement Learning for Mobile Manipulation
- Deep Reinforcement Learning for Active High Frequency Trading
- Autonomous UAV Base Stations for Next Generation Wireless Networks: A Deep Learning Approach
- Enhancing Efficiency and Propulsion in Bio-mimetic Robotic Fish through End-to-End Deep Reinforcement Learning
- Green Deep Reinforcement Learning for Radio Resource Management: Architecture, Algorithm Compression and Challenge
- Flexible and Efficient Long-Range Planning Through Curious Exploration
- A Continual Offline Reinforcement Learning Benchmark for Navigation Tasks
- An Introduction to Deep Reinforcement Learning
- Are Gradient-based Saliency Maps Useful in Deep Reinforcement Learning?
- Transfer learning-enhanced deep reinforcement learning for aerodynamic airfoil optimisation subject to structural constraints
- Explainable Artificial Intelligence (XAI) for 6G: Improving Trust between Human and Machine
- A Survey of Behavior Learning Applications in Robotics -- State of the Art and Perspectives
- A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law
- BF++: a language for general-purpose program synthesis
- Techniques Toward Optimizing Viewability in RTB Ad Campaigns Using Reinforcement Learning
- Is Deep Reinforcement Learning Ready for Practical Applications in Healthcare? A Sensitivity Analysis of Duel-DDQN for Hemodynamic Management in Sepsis Patients
- Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement Learning
- Deep Learning in Physical Layer Communications
- Learning to be Global Optimizer
- Navigate the Unknown: Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration
- LLM-Powered AI Agent Systems and Their Applications in Industry
- Remote Sensing Object Tracking With Deep Reinforcement Learning Under Occlusion
- Diffmv: A Unified Diffusion Framework for Healthcare Predictions with Random Missing Views and View Laziness
- Parallel Placement of Virtualized Network Functions via Federated Deep Reinforcement Learning
- Dynamics of striatal action selection and reinforcement learning
- Model reusability in Reinforcement Learning
- Design and Experimental Test of Datatic Approximate Optimal Filter in Nonlinear Dynamic Systems
- Deep Learning: A Comprehensive Overview on Techniques, Taxonomy, Applications and Research Directions
- A Generalised and Adaptable Reinforcement Learning Stopping Method
- Fuzzy–TD3: A prior-knowledge guided deterministic policy gradient for interpretable control
- Neural Network-Based Uncertainty Quantification: A Survey of Methodologies and Applications
- MimicBot: Combining Imitation and Reinforcement Learning to win in Bot Bowl
- Reinforcement Learning with Convolutional Reservoir Computing
- Quantum-Enhanced Hybrid Reinforcement Learning Framework for Dynamic Path Planning in Autonomous Systems
- Motion Generation for Food Topping Challenge 2024: Serving Salmon Roe Bowl and Picking Fried Chicken
- BQSched: A Non-intrusive Scheduler for Batch Concurrent Queries via Reinforcement Learning
- Physical Layer Security for Integrated Sensing and Communication: A Survey
- Stigmergic Independent Reinforcement Learning for Multi-Agent Collaboration
- Efficient Tree Generation for Globally Optimal Decisions under Probabilistic Outcomes
- A Systematic Approach to Design Real-World Human-in-the-Loop Deep Reinforcement Learning: Salient Features, Challenges and Trade-offs
- Reinforcement learning for suppression of collective activity in oscillatory ensembles
- Mathematical methods of reinforcement learning
- Deep Learning -- A first Meta-Survey of selected Reviews across Scientific Disciplines, their Commonalities, Challenges and Research Impact
- A Deep Reinforcement Learning Approach for the Meal Delivery Problem
- Communication-Computation Efficient Device-Edge Co-Inference via AutoML
- Diversity-based Trajectory and Goal Selection with Hindsight Experience Replay
- AI in a vat: Fundamental limits of efficient world modelling for agent sandboxing and interpretability
- Proportional integral derivative controller assisted reinforcement learning for path following by autonomous underwater vehicles
- Improving RL Exploration for LLM Reasoning through Retrospective Replay
- Practical Deep Learning Architecture Optimization
- Learning to Selectively Transfer
- The wireless control plane: An overview and directions for future research
- BenchQC -- Scalable and modular benchmarking of industrial quantum computing applications
- Moderate Actor-Critic Methods: Controlling Overestimation Bias via Expectile Loss
Discussions
Related