Playing Atari with Deep Reinforcement Learning
2013/12/19 by Volodymyr Mnih, Koray Kavukcuoglu, Mnih, Volodymyr +11 · 5 voices · 361 citations
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Artificial Intelligence in Games #Reinforcement Learning in Robotics #cs.LG
paper · pdf · doi:10.48550/arxiv.1312.5602
openalex publication_date 2013/12/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Abstract
We present the first deep learning model to successfully learn control policies directly from high-dimensional sensory input using reinforcement learning. The model is a convolutional neural network, trained with a variant of Q-learning, whose input is raw pixels and whose output is a value function estimating future rewards. We apply our method to seven Atari 2600 games from the Arcade Learning Environment, with no adjustment of the architecture or learning algorithm. We find that it outperforms all previous approaches on six of the games and surpasses a human expert on three of them.
Citations
Cited by
- Error Amplification Limits ANN-to-SNN Conversion in Continuous Control
- What Matters for Simulation to Online Reinforcement Learning on Real Robots
- From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning
- Gaussian-Mixture-Model Q-Functions for Policy Iteration in Reinforcement Learning
- HypEMBER: Hypernetwork-based Ensemble for Robust Policy Learning of Parametrized Dynamical Systems
- Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems
- Co-Design of Aeroelastic Systems with Deep Reinforcement Learning
- Memory Merge DQN: Sensitivity Weighted Target Updates for Stable Value Learning
- Molecular quantum control algorithm design by reinforcement learning
- From Particles to Perils: SVGD-Based Hazardous Scenario Generation for Autonomous Driving Systems Testing
- Partial recurrence enables robust and efficient computation
- Hierarchical Reasoning Model
- Reinforcement Learning Teachers of Test Time Scaling
- Can machines learn density functionals? Past, present, and future of ML in DFT
- Evolution and The Knightian Blindspot of Machine Learning
- The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
- Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement Learning
- Nash Equilibrium Between Consumer Electronic Devices and DoS Attacker for Distributed IoT-enabled RSE Systems
- VulnGym: Evaluating Vulnerability Management Strategies against Advanced Persistent Threats
- CAST: Game Solvers as Turn-Level Teachers for LLM Agents
- Calibrated Partial Resets: Preventing Policy Collapse in Continual Reinforcement Learning
- Endpoint Replay: Compressing the Recency Buffer in Deep Reinforcement Learning
- A Survey of Freshness-Aware Wireless Networking with Reinforcement Learning
- Recurrent Off-Policy Deep Reinforcement Learning Doesn't Have to be Slow
- LacaDM: A Latent Causal Diffusion Model for Multiobjective Reinforcement Learning
- Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
- CoDrone: Autonomous Drone Navigation Assisted by Edge and Cloud Foundation Models
- Online Robust Reinforcement Learning with General Function Approximation
- Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
- Demonstration-Guided Continual Reinforcement Learning in Dynamic Environments
- StarCraft+: Benchmarking Multi-agent Algorithms in Adversary Paradigm
- VET Your Agent: Towards Host-Independent Autonomy via Verifiable Execution Traces
- Spectral Representation-based Reinforcement Learning
- Tiny, On-Device Decision Makers with the MiniConv Library
- FM-EAC: Feature Model-based Enhanced Actor-Critic for Multi-Task Control in Dynamic Environments
- Goal Reaching with Eikonal-Constrained Hierarchical Quasimetric Reinforcement Learning
- Adaptive Replay Buffer for Offline-to-Online Reinforcement Learning
- Push Smarter, Not Harder: Hierarchical RL-Diffusion Policy for Efficient Nonprehensile Manipulation
- Advancing LLM-Based Security Automation with Customized Group Relative Policy Optimization for Zero-Touch Networks
- TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
- Scalable Offline Model-Based RL with Action Chunks
- Deep Reinforcement Learning for Phishing Detection with Transformer-Based Semantic Features
- A Fast Anti-Jamming Cognitive Radar Deployment Algorithm Based on Reinforcement Learning
- Marti-5: A Mathematical Model of "Self in the World" as a First Step Toward Self-Awareness
- AI & Human Co-Improvement for Safer Co-Superintelligence
- Using Machine Learning to Take Stay-or-Go Decisions in Data-driven Drone Missions
- Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-Function
- Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism
- Deep Reinforcement Learning for Dynamic Algorithm Configuration: A Case Study on Optimizing OneMax with the (1+(λ,λ))-GA
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- TECM*: A Data-Driven Assessment to Reinforcement Learning Methods and Application to Heparin Treatment Strategy for Surgical Sepsis
- Dynamic Configuration of On-Street Parking Spaces using Multi Agent Reinforcement Learning
- Differentiable Weightless Controllers: Learning Logic Circuits for Continuous Control
- Extending NGU to Multi-Agent RL: A Preliminary Study
- Partially Equivariant Reinforcement Learning in Symmetry-Breaking Environments
- Hardware-Software Collaborative Computing of Photonic Spiking Reinforcement Learning for Robotic Continuous Control
- SmallWorlds: Assessing Dynamics Understanding of World Models in Isolated Environments
- Heterogeneous Multi-Agent Reinforcement Learning with Attention for Cooperative and Scalable Feature Transformation
- Distributed quantum architecture search using multi-agent reinforcement learning
- Adaptive Dueling Double Deep Q-networks in Uniswap V3 Replication and Extension with Mamba
- Attention Trajectories as a Diagnostic Axis for Deep Reinforcement Learning
- InF-ATPG: Intelligent FFR-Driven ATPG with Advanced Circuit Representation Guided Reinforcement Learning
- Video Object Recognition in Mobile Edge Networks: Local Tracking or Edge Detection?
- SpeedAug: Policy Acceleration via Tempo-Enriched Policy and RL Fine-Tuning
- MOMA-AC: A preference-driven actor-critic framework for continuous multi-objective multi-agent reinforcement learning
- Physical Reinforcement Learning
- LAOF: Robust Latent Action Learning with Optical Flow Constraints
- Limitations of Scalarisation in MORL: A Comparative Study in Discrete Environments
- IPR-1: Interactive Physical Reasoner
- Simulated Human Learning in a Dynamic, Partially-Observed, Time-Series Environment
- Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning
- Object-Centric World Models for Causality-Aware Reinforcement Learning
- Boosting Reinforcement Learning in 3D Visuospatial Tasks Through Human-Informed Curriculum Design
- NFQ2.0: The CartPole Benchmark Revisited
- PROF: An LLM-based Reward Code Preference Optimization Framework for Offline Imitation Learning
- Graph Neural Networks, Deep Reinforcement Learning and Probabilistic Topic Modeling for Strategic Multiagent Settings
- Simulator and Experience Enhanced Diffusion Model for Comprehensive ECG Generation
- Representation Learning Enables Scalable Multitask Deep Reinforcement Learning
- Multistep Quasimetric Learning for Scalable Goal-conditioned Reinforcement Learning
- Enabling Off-Policy Imitation Learning with Deep Actor Critic Stabilization
- Source-Only Cross-Weather LiDAR via Geometry-Aware Point Drop
- Approximating Shapley Explanations in Reinforcement Learning
- Deep Reinforcement Learning for Dynamic Origin-Destination Matrix Estimation in Microscopic Traffic Simulations Considering Credit Assignment
- Multi-agent Coordination via Flow Matching
- Environment Agnostic Goal-Conditioning, A Study of Reward-Free Autonomous Learning
- Enhancing Q-Value Updates in Deep Q-Learning via Successor-State Prediction
- DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay
- From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement Learning
- None To Optima in Few Shots: Bayesian Optimization with MDP Priors
- Power Control Based on Multi-Agent Deep Q Network for D2D Communication
- Challenges in Credit Assignment for Multi-Agent Reinforcement Learning in Open Agent Systems
- Towards Reinforcement Learning Based Log Loading Automation
- Conformal Prediction Beyond the Horizon: Distribution-Free Inference for Policy Evaluation
- To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning
- Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning
- A Benchmark Study of Deep Reinforcement Learning Algorithms for the Container Stowage Planning Problem
- Expanding LLM Agent Boundaries with Strategy-Guided Exploration
- Improving Heart-Focused Medical Question Answering in LLMs via Variance-Aware Rubric Rewards with GRPO
- NavQ: Learning a Q-Model for Foresighted Vision-and-Language Navigation
- Dense and Diverse Goal Coverage in Multi Goal Reinforcement Learning
- Learning-Based vs Human-Derived Congestion Control: An In-Depth Experimental Study
- Adaptive Surrogate Gradients for Sequential Reinforcement Learning in Spiking Neural Networks
- Survey and Tutorial of Reinforcement Learning Methods in Process Systems Engineering
- Hybrid Modeling, Sim-to-Real Reinforcement Learning, and Large Language Model Driven Control for Digital Twins
- Transitive RL: Value Learning via Divide and Conquer
- Reinforcement learning-guided optimization of critical current in high-temperature superconductors
- Is Temporal Difference Learning the Gold Standard for Stitching in RL?
- Narrowing Action Choices with AI Improves Human Sequential Decisions
- Software Engineering Agents for Embodied Controller Generation : A Study in Minigrid Environments
- PARL: Prompt-based Agents for Reinforcement Learning
- ProSh: Probabilistic Shielding for Model-free Reinforcement Learning
- Balancing Specialization and Centralization: A Multi-Agent Reinforcement Learning Benchmark for Sequential Industrial Control
- COLA: Continual Learning via Autoencoder Retrieval of Adapters
- StarBench: A Turn-Based RPG Benchmark for Agentic Multimodal Decision-Making and Information Seeking
- Actor-Free Continuous Control via Structurally Maximizable Q-Functions
- REPAIR Approach for Social-based City Reconstruction Planning in case of natural disasters
- Heterogeneous Adversarial Play in Interactive Environments
- Learning to Design Soft Hands using Reward Models
- Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time Control
- Learning to play: A Multimodal Agent for 3D Game-Play
- Human-Allied Relational Reinforcement Learning
- Zero-shot World Models via Search in Memory
- The Pursuit of Diversity: Multi-Objective Testing of Deep Reinforcement Learning Agents
- Stop-RAG: Value-Based Retrieval Control for Iterative RAG
- ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning
- Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
- Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
- AOAD-MAT: Transformer-based multi-agent deep reinforcement learning model considering agents' order of action decisions
- Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
- MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
- Neutral Agent-based Adversarial Policy Learning against Deep Reinforcement Learning in Multi-party Open Systems
- ExGRPO: Learning to Reason from Experience
- Efficient Restarts in Non-Stationary Model-Free Reinforcement Learning
- High-Dimensional Learning Dynamics of Quantized Models with Straight-Through Estimator
- PAC-Bayesian Reinforcement Learning Trains Generalizable Policies
- Experience-Efficient Model-Free Deep Reinforcement Learning Using Pre-Training
- Agent Learning via Early Experience
- Energy-Guided Diffusion Sampling for Long-Term User Behavior Prediction in Reinforcement Learning-based Recommendation
- Rethinking Provenance Completeness with a Learning-Based Linux Scheduler
- Dual Goal Representations
- Local Reinforcement Learning with Action-Conditioned Root Mean Squared Q-Functions
- Medical Vision Language Models as Policies for Robotic Surgery
- AutoPentester: An LLM Agent-based Framework for Automated Pentesting
- BuilderBench -- A benchmark for generalist agents
- Trade in Minutes! Rationality-Driven Agentic System for Quantitative Financial Trading
- Wavelet Predictive Representations for Non-Stationary Reinforcement Learning
- Q-Learning with Shift-Aware Upper Confidence Bound in Non-Stationary Reinforcement Learning
- Deep Reinforcement Learning for Multi-Agent Coordination
- Memory-Driven Self-Improvement for Decision Making with Large Language Models
- RL-Guided Data Selection for Language Model Finetuning
- Learning Distinguishable Representations in Deep Q-Networks for Linear Transfer
- Quantifying Generalisation in Imitation Learning
- Beyond Softmax: A Natural Parameterization for Categorical Random Variables
- Optimisation of Resource Allocation in Heterogeneous Wireless Networks Using Deep Reinforcement Learning
- STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning
- CaRe-BN: Precise Moving Statistics for Stabilizing Spiking Neural Networks in Reinforcement Learning
- Bridging Discrete and Continuous RL: Stable Deterministic Policy Gradient with Martingale Characterization
- Learning Market Making with Closing Auctions
- Reinforcement Learning-Based Prompt Template Stealing for Text-to-Image Models
- Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
- Activation Function Design Sustains Plasticity in Continual Learning
- Impact of Collective Behaviors of Autonomous Vehicles on Urban Traffic Dynamics: A Multi-Agent Reinforcement Learning Approach
- ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
- Physics of Learning: A Lagrangian perspective to different learning paradigms
- Adaptive Federated Learning Defences via Trust-Aware Deep Q-Networks
- Sig2Model: A Boosting-Driven Model for Updatable Learned Indexes
- Inverse Reinforcement Learning with Just Classification and a Few Regressions
- Wonder Wins Ways: Curiosity-Driven Exploration through Multi-Agent Contextual Calibration
- Leveraging Temporally Extended Behavior Sharing for Multi-task Reinforcement Learning
- Embodied AI: From LLMs to World Models
- Understanding Adversarial Attacks on Observations in Deep Reinforcement Learning
- PrintAnything: Learning an Intermediate Representation for 3D printing G-code Generation
- AutoSoC: Automating Algorithm-SOC Co-design for Aerial Robots
- SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
- Moving by Looking: Towards Vision-Driven Avatar Motion Generation
- SlicePilot: Demystifying Network Slice Placement in Heterogeneous Cloud Infrastructures
- Efficient Reinforcement Learning by Reducing Forgetting with Elephant Activation Functions
- CayleyPy Growth: Efficient growth computations and hundreds of new conjectures on Cayley graphs (Brief version)
- Learning View and Target Invariant Visual Servoing for Navigation
- Towards General Computer Control with Hierarchical Agents and Multi-Level Action Spaces
- On the Limits of Tabular Hardness Metrics for Deep RL: A Study with the Pharos Benchmark
- Learning from Observation: A Survey of Recent Advances
- The Distribution Shift Problem in Transportation Networks using Reinforcement Learning and AI
- Online reinforcement learning via sparse Gaussian mixture model Q-functions
- TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
- Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
- Zero-sum turn games using Q-learning: finite computation with security guarantees
- PuzzleJAX: A Benchmark for Reasoning and Learning
- Linear Dynamics meets Linear MDPs: Closed-Form Optimal Policies via Reinforcement Learning
- Compositional shield synthesis for safe reinforcement learning in partial observability
- A Practical Adversarial Attack against Sequence-based Deep Learning Malware Classifiers
- Gradient Free Deep Reinforcement Learning With TabPFN
- A neural drift-plus-penalty algorithm for network power allocation and routing
- How well can LLMs provide planning feedback in grounded environments?
- Reinforcement Learning based Dynamic Model Selection for Short-Term Load Forecasting
- One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning
- Language Self-Play For Data-Free Training
- Timing the Message: Language-Based Notifications for Time-Critical Assistive Settings
- Physics-informed Value Learner for Offline Goal-Conditioned Reinforcement Learning
- TrajAware: Graph Cross-Attention and Trajectory-Aware for Generalisable VANETs under Partial Observations
- FinXplore: An Adaptive Deep Reinforcement Learning Framework for Balancing and Discovering Investment Opportunities
- An Arbitration Control for an Ensemble of Diversified DQN variants in Continual Reinforcement Learning
- Improving Robustness of AlphaZero Algorithms to Test-Time Environment Changes
- Reinforced Data-Driven Estimation for Spectral Properties of Koopman Semigroup in Stochastic Dynamical Systems
- Off-Policy Actor-Critic with Shared Experience Replay
- Deep Reinforcement Learning for Sponsored Search Real-time Bidding
- Multi-Agent Reinforcement Learning for Task Offloading in Wireless Edge Networks
- Machine Intelligence on the Edge: Interpretable Cardiac Pattern Localisation Using Reinforcement Learning
- Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning
- Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models
- ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition
- Breaking the Cold-Start Barrier: Reinforcement Learning with Double and Dueling DQNs
- Improving Aviation Safety Analysis: Automated HFACS Classification Using Reinforcement Learning with Group Relative Policy Optimization
- Off-Policy Self-Critical Training for Transformer in Visual Paragraph Generation
- Out-of-Distribution Dynamics Detection: RL-Relevant Benchmarks and\n Results
- Machine-Learning-Assisted Pulse Design for State Preparation in a Noisy Environment
- The Effectiveness of Memory Replay in Large Scale Continual Learning
- Demystifying Reward Design in Reinforcement Learning for Upper Extremity Interaction: Practical Guidelines for Biomechanical Simulations in HCI
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition
- Online Hyper-parameter Tuning in Off-policy Learning via Evolutionary Strategies
- Optimistic Distributionally Robust Policy Optimization
- Financial Decision Making using Reinforcement Learning with Dirichlet Priors and Quantum-Inspired Genetic Optimization
- Real-time visual tracking by deep reinforced decision making
- Learning Game-Playing Agents with Generative Code Optimization
- PoolFlip: A Multi-Agent Reinforcement Learning Security Environment for Cyber Defense
- Deep Reinforcement Learning for Complex Manipulation Tasks with Sparse Feedback
- Real-Time Model Checking for Closed-Loop Robot Reactive Planning
- Skill-Aligned Fairness in Multi-Agent Learning for Collaboration in Healthcare
- Language Models For Generalised PDDL Planning: Synthesising Sound and Programmatic Policies
- Uncertainty-Aware Reinforcement Learning for Collision Avoidance
- Adversarial Agent Behavior Learning in Autonomous Driving Using Deep Reinforcement Learning
- A Dynamical Systems Framework for Reinforcement Learning Safety and Robustness Verification
- A Reinforcement Learning Approach to the View Planning Problem
- Overcoming Digital Gravity when using AI in Public Health Decisions
- Slipping to the Extreme: A Mixed Method to Explain How Extreme Opinions Infiltrate Online Discussions
- Stability-certified reinforcement learning: A control-theoretic perspective
- SDN Flow Entry Management Using Reinforcement Learning
- Beyond ReLU: Chebyshev-DQN for Enhanced Deep Q-Networks
- Deep Reinforcement Learning for Imbalanced Classification
- Large Batch Training Does Not Need Warmup
- Pixels to Play: A Foundation Model for 3D Gameplay
- Competitive Experience Replay
- A Short Survey On Memory Based Reinforcement Learning
- Soft Actor-Critic Algorithms and Applications
- Adversarial recovery of agent rewards from latent spaces of the limit order book
- PAPPL: Personalized AI-Powered Progressive Learning Platform
- MDPO: Overcoming the Training-Inference Divide of Masked Diffusion Language Models
- Boosting Image Recognition with Non-differentiable Constraints
- Discovering hierarchies using Imitation Learning from hierarchy aware policies
- Predictor-Corrector Policy Optimization
- Towards Content Provider Aware Recommender Systems: A Simulation Study on the Interplay between User and Provider Utilities
- Data-Efficient Reinforcement Learning for Malaria Control
- Task-Agnostic Morphology Evolution
- MAPEL: Multi-Agent Pursuer-Evader Learning using Situation Report
- Meta-Learning Bandit Policies by Gradient Ascent
- A Comprehensive Review of AI Agents: Transforming Possibilities in Technology and Beyond
- Approximate Inference with Amortised MCMC
- D2Q Synchronizer: Distributed SDN Synchronization for Time Sensitive Applications
- Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework
- Developing a Simple Model for Sand-Tool Interaction and Autonomously Shaping Sand
- Understanding in Artificial Intelligence
- Optimizing Large-Scale Fleet Management on a Road Network using Multi-Agent Deep Reinforcement Learning with Graph Neural Network
- Hierarchically Integrated Models: Learning to Navigate from Heterogeneous Robots
- RUDDER: Return Decomposition for Delayed Rewards
- Reinforced Language Models for Sequential Decision Making
- Reinforcement-Learning-Designed Field-Free Sub-Nanosecond Spin-Orbit-Torque Switching
- Offline Reinforcement Learning with Soft Behavior Regularization
- Improve Agents without Retraining: Parallel Tree Search with Off-Policy Correction
- Dirichlet Pruning for Neural Network Compression
- Playing Catan with Cross-dimensional Neural Network
- Multi-Agent Path Planning based on MPC and DDPG
- Distributional Reinforcement Learning for Multi-Dimensional Reward\n Functions
- A Review On Safe Reinforcement Learning Using Lyapunov and Barrier Functions
- Continual Learning in Neural Networks
- SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokens
- LPaintB: Learning to Paint from Self-Supervision
- Toward Lifelong Learning in Equilibrium Propagation: Sleep-like and Awake Rehearsal for Enhanced Stability
- Playing Atari Space Invaders with Sparse Cosine Optimized Policy Evolution
- An Intelligent Control Strategy for buck DC-DC Converter via Deep Reinforcement Learning
- Meta Reinforcement Learning with Distribution of Exploration Parameters Learned by Evolution Strategies
- ReconDreamer-RL: Enhancing Reinforcement Learning via Diffusion-based Scene Reconstruction
- From proprioception to long-horizon planning in novel environments: A\n hierarchical RL model
- Reinforcement learning for batch bioprocess optimization
- Evolutionary Optimization of Deep Learning Agents for Sparrow Mahjong
- A State Aggregation Approach for Solving Knapsack Problem with Deep Reinforcement Learning
- Efficient Ridesharing Dispatch Using Multi-Agent Reinforcement Learning
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- PANAMA: A Network-Aware MARL Framework for Multi-Agent Path Finding in Digital Twin Ecosystems
- Progressive Relation Learning for Group Activity Recognition
- Multimedia Edge Computing
- Robusta: Robust AutoML for Feature Selection via Reinforcement Learning
- Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees
- CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL
- Smaller Models, Better Generalization
- Multimodal LLM-assisted Evolutionary Search for Programmatic Control Policies
- HCRide: Harmonizing Passenger Fairness and Driver Preference for Human-Centered Ride-Hailing
- Deep reinforcement learning to detect brain lesions on MRI: a\n proof-of-concept application of reinforcement learning to medical images
- Model-based Adversarial Meta-Reinforcement Learning
- DiWA: Diffusion Policy Adaptation with World Models
- Fast reinforcement learning for decentralized MAC optimization
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- A Non-Technical Survey on Deep Convolutional Neural Network Architectures
- Replay in Deep Learning: Current Approaches and Missing Biological Elements
- Reinforcement Learning with General Value Function Approximation: Provably Efficient Approach via Bounded Eluder Dimension
- Transfer Learning Across Patient Variations with Hidden Parameter Markov Decision Processes
- CO-RFT: Efficient Fine-Tuning of Vision-Language-Action Models through Chunked Offline Reinforcement Learning
- Recurrent Value Functions
- ACTRCE: Augmenting Experience via Teacher's Advice For Multi-Goal Reinforcement Learning
- A Machine Learning Approach to Routing
- MRPB 1.0: A Unified Benchmark for the Evaluation of Mobile Robot Local Planning Approaches
- Learning Network Dismantling Without Handcrafted Inputs
- Task-Relevant Object Discovery and Categorization for Playing First-person Shooter Games
- Hierarchical Message-Passing Policies for Multi-Agent Reinforcement Learning
- Collaborative creativity with Monte-Carlo Tree Search and Convolutional Neural Networks
- Safe Reinforcement Learning Using Robust Action Governor
- Cooperative multi-agent reinforcement learning for high-dimensional\n nonequilibrium control
- Personalized Education with Ranking Alignment Recommendation
- Dynamics-Aware Unsupervised Discovery of Skills
- Benchmarking Massively Parallelized Multi-Task Reinforcement Learning for Robotics Tasks
- Deep Implicit Coordination Graphs for Multi-agent Reinforcement Learning
- Enhancing Multi-Agent Collaboration with Attention-Based Actor-Critic Policies
- Boosting Soft Actor-Critic: Emphasizing Recent Experience without Forgetting the Past
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform Sampling
- PreCo: A Large-scale Dataset in Preschool Vocabulary for Coreference Resolution
- DeepGo: Predictive Directed Greybox Fuzzing
- What Does it Mean for a Neural Network to Learn a "World Model"?
- MoB Queen: Mixture of Bandits with a Global Density Map
- Giraffe: Using Deep Reinforcement Learning to Play Chess
- Tracking the Race Between Deep Reinforcement Learning and Imitation Learning -- Extended Version
- FAST: Similarity-based Knowledge Transfer for Efficient Policy Learning
- Evolutionary Population Curriculum for Scaling Multi-Agent Reinforcement Learning
- Machine Learning With Neuromorphic Photonics
- Influencing Towards Stable Multi-Agent Interactions
- Region Growing Curriculum Generation for Reinforcement Learning
- Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection
- Implicit Generation and Generalization in Energy-Based Models
- Causal Confusion in Imitation Learning
- Minding Motivation: The Effect of Intrinsic Motivation on Agent Behaviors
- Hierarchical Reinforcement Learning Method for Autonomous Vehicle Behavior Planning
- Quantum Reinforcement Learning by Adaptive Non-local Observables
- Hierarchical Deep Reinforcement Learning Framework for Multi-Year Asset Management Under Budget Constraints
- Virne: A Comprehensive Benchmark for Deep RL-based Network Resource Allocation in NFV
- Recurrent Predictive State Policy Networks
- Multi-Year Maintenance Planning for Large-Scale Infrastructure Systems: A Novel Network Deep Q-Learning Approach
- Reinforcement Learning via Conservative Agent for Environments with Random Delays
- The Dynamics of Handwriting Improves the Automated Diagnosis of Dysgraphia
- SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning
- Learning Visual Predictive Models of Physics for Playing Billiards
- OPAC: Opportunistic Actor-Critic
- Deep Reinforcement Learning-based Task Offloading in Satellite-Terrestrial Edge Computing Networks
- ADARES: Adaptive Resource Management for Virtual Machines
- Uniform State Abstraction For Reinforcement Learning
- AutoPilot: Automating SoC Design Space Exploration for SWaP Constrained\n Autonomous UAVs
- A Dual-Hormone Closed-Loop Delivery System for Type 1 Diabetes Using Deep Reinforcement Learning
- Vizarel: A System to Help Better Understand RL Agents
- Adaptive perturbation adversarial training: based on reinforcement learning
- Self-Tuning Sectorization: Deep Reinforcement Learning Meets Broadcast Beam Optimization
- GASL: Guided Attention for Sparsity Learning in Deep Neural Networks
- Deep Reinforcement Learning for Tactile Robotics: Learning to Type on a Braille Keyboard
- Mapping Instructions to Actions in 3D Environments with Visual Goal\n Prediction
- Randomized Policy Learning for Continuous State and Action MDPs
Discussions
- Playing Atari with Deep Reinforcement Learning [pdf] [hn, 37 points, 9 comments]
- Playing Atari with Deep Reinforcement Learning (2013) [pdf] [hn, 21 points, 2 comments]
- Deep Mind's neural networks play 7 Atari 2600 games with more skill than a human [hn, 5 points, 0 comments]
- Playing Atari with Deep Reinforcement Learning [hn, 3 points, 1 comments]
- And the paper that started the trend (way back in 2013): arxiv.org/abs/1312.5602 [bsky, 1 points, 0 comments]
Related