Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
1999/08/01 by Richard S. Sutton, Doina Precup, Satinder Singh · 3,169 citations
Computer Science · Mathematics · #Abstraction #Action (physics) #Advanced Software Engineering Methodologies #Artificial intelligence #Computer science #Dynamic programming #Formal Methods in Verification #Machine learning #Markov decision process #Markov process #Mathematics #Process (computing) #Programming language #Reinforcement Learning in Robotics #Reinforcement learning #Set (abstract data type)
paper · pdf · doi:10.1016/s0004-3702(99)00052-1
published in Artificial Intelligence 112(1-2), 181-211 (Elsevier BV)
openalex publication_date 1999/08/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04
Citations
Cited by
- Beyond Fine-Tuning: Transferring Behavior in Reinforcement Learning
- Contextualizing predictive minds
- Exploring the hierarchical structure of human plans via program generation
- When Is a Learned Command Adapter Worth It? Closed-Loop Identification and Counterfactual Auditing of Frozen Locomotion Policies
- Learning, Reward, and Decision Making
- Training Fast Robot Policies with Slow Foundation Models
- HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents
- S3: Stable Subgoal Selection by Constraining Uncertainty of Coarse Dynamics in Hierarchical Reinforcement Learning
- Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
- SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning
- Retriever: Composing Closed-Loop Asynchronous Robot Programs
- The challenge of hidden gifts in multi-agent reinforcement learning
- PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation
- Counterfactual Shapley Credit Assignment
- Orbis 2: A Hierarchical World Model for Driving
- When to Plan: Learning to Select Between Reactive Control and Deliberative Planning
- Foresight Residual RL for Long-Horizon Robot Manipulation with Vision-Language-Action Models
- The Terminal Representation in Reinforcement Learning
- Structure-Induced Information for Rerooting Levin Tree Search
- A Compositional Framework for Open-ended Intelligence
- Separable Pathways for Causal Reasoning: How Architectural Scaffolding Enables Hypothesis-Space Restructuring in LLM Agents
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Learning Abstractions for Hierarchical Planning in Program-Synthesis Agents
- Deep Reinforcement Learning Based Navigation with Macro Actions and Topological Maps
- Modeling Others' Minds as Code
- Reward-Aware Proto-Representations in Reinforcement Learning
- Synthesizing world models for bilevel planning
- Dr Jekyll and Mr Hyde: the Strange Case of Off-Policy Policy Updates
- Closing the gap towards end-to-end autonomous vehicle system
- Equilibrium Inverse Reinforcement Learning for Ride-hailing Vehicle Network
- Adversarial Option-Aware Hierarchical Imitation Learning
- Sub-Goal Trees -- a Framework for Goal-Based Reinforcement Learning
- RLocator: Reinforcement Learning for Bug Localization
- Efficient Deep Reinforcement Learning via Adaptive Policy Transfer
- Reinforcement Networks: novel framework for collaborative Multi-Agent Reinforcement Learning tasks
- Deep Successor Reinforcement Learning
- Learning One Representation to Optimize All Rewards
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
- Disentangled Skill Embeddings for Reinforcement Learning
- DialPort: Connecting the Spoken Dialog Research Community to Real User Data
- Attention Option-Critic
- Beyond Tabula-Rasa: a Modular Reinforcement Learning Approach for Physically Embedded 3D Sokoban
- RODE: Learning Roles to Decompose Multi-Agent Tasks
- An operator view of policy gradient methods
- Value Prediction Network
- Ask Your Humans: Using Human Instructions to Improve Generalization in Reinforcement Learning
- Compositional Behavioral Semantics for State Abstraction in Reinforcement Learning
- Active Learning for Autonomous Intelligent Agents: Exploration, Curiosity, and Interaction
- Learning High-level Representations from Demonstrations
- Hierarchical Soft Actor-Critic for Sparse-Reward Long-Horizon Reinforcement Learning
- WCM: World-Cognition Model for Generalizable Human-Robot Interaction
- Interpretable Model-based Hierarchical Reinforcement Learning using Inductive Logic Programming
- The Game of Tetris in Machine Learning
- MEMENTO: Memory-Guided Memetic Code-as-Policy Evolution
- Options of Interest: Temporal Abstraction with Interest Functions
- DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification
- HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
- Hierarchical Subtask Discovery With Non-Negative Matrix Factorization
- Learning Skills from Action-Free Videos
- Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
- Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
- Monte Carlo Query Search: Active Capability Assessment of AI Agents
- Model-Based Reinforcement Learning in Discrete-Action Non-Markovian Reward Decision Processes
- Prediction and Control with Temporal Segment Models
- Sample-Efficient Robot Skill Learning for Construction Tasks: Benchmarking Hierarchical Reinforcement Learning and Vision-Language-Action VLA Model
- Learning Abstract Options
- A Definition of Open-Ended Learning Problems for Goal-Conditioned Agents
- CoDA: A Context-Decoupled Hierarchical Agent with Reinforcement Learning
- Goal Reaching with Eikonal-Constrained Hierarchical Quasimetric Reinforcement Learning
- Towards Trustworthy Multi-Turn LLM Agents via Behavioral Guidance
- Learnings Options End-to-End for Continuous Action Tasks
- A First-Occupancy Representation for Reinforcement Learning
- Push Smarter, Not Harder: Hierarchical RL-Diffusion Policy for Efficient Nonprehensile Manipulation
- SCOPE: Language Models as One-Time Teacher for Hierarchical Planning in Text Environments
- A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms
- On Learning to Think: Algorithmic Information Theory for Novel\n Combinations of Reinforcement Learning Controllers and Recurrent Neural World\n Models
- Disentangling causal effects for hierarchical reinforcement learning
- Variational Intrinsic Control
- Opponent interactions between serotonin and dopamine
- Probabilistic Mixture-of-Experts for Efficient Deep Reinforcement Learning
- Learning When to Switch: Adaptive Policy Selection via Reinforcement Learning
- Realizable Abstractions: Near-Optimal Hierarchical Reinforcement Learning
- Task diversity produces systematic transfer but inhibits continual reinforcement learning
- Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
- Autonomous Penetration Testing using Reinforcement Learning
- How Does RL Post-training Induce Skill Composition? A Case Study on Countdown
- Deep Learning in Neural Networks: An Overview
- Constant-Time Motion Planning with Manipulation Behaviors
- Language as an Abstraction for Hierarchical Deep Reinforcement Learning
- ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
- Safe and Sustainable Electric Bus Charging Scheduling with Constrained Hierarchical DRL
- Logically-Constrained Reinforcement Learning
- Algorithms for Batch Hierarchical Reinforcement Learning
- Learning to Generalize Across Long-Horizon Tasks from Human Demonstrations
- A New Error Temporal Difference Algorithm for Deep Reinforcement Learning in Microgrid Optimization
- Noise-Adaptive Quantum Circuit Mapping for Multi-Chip NISQ Systems via Deep Reinforcement Learning
- Convergence and stability of Q-learning in Hierarchical Reinforcement Learning
- Multi-Level Discovery of Deep Options
- Retrospective Analysis of the 2019 MineRL Competition on Sample Efficient Reinforcement Learning
- Conditional Diffusion Model for Multi-Agent Dynamic Task Decomposition
- Learning to Interrupt: A Hierarchical Deep Reinforcement Learning Framework for Efficient Exploration
- Neural Programmer-Interpreters
- Deep Reinforcement Learning From Raw Pixels in Doom
- Momentum Centering and Asynchronous Update for Adaptive Gradient Methods
- Deciding What to Learn: A Rate-Distortion Approach
- Leveraging Human Guidance for Deep Reinforcement Learning Tasks
- Context-Aware Policy Reuse
- Independently Controllable Features
- Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied Navigation
- Long Text Generation via Adversarial Training with Leaked Information
- Incorporating Spatial Information into Goal-Conditioned Hierarchical Reinforcement Learning via Graph Representations
- Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue System
- Self-Abstraction from Grounded Experience for Plan-Guided Policy Refinement
- Monitor-Generate-Verify (MGV): Formalising Metacognitive Theory for Language Model Reasoning
- Multitasking Inhibits Semantic Drift
- Learning Interactive World Model for Object-Centric Reinforcement Learning
- SLAP: Shortcut Learning for Abstract Planning
- Reinforcement Learning for Pollution Detection in a Randomized, Sparse and Nonstationary Environment with an Autonomous Underwater Vehicle
- Learning to Plan & Schedule with Reinforcement-Learned Bimanual Robot Skills
- Skew-Fit: State-Covering Self-Supervised Reinforcement Learning
- Decoupled Learning of Environment Characteristics for Safe Exploration
- Composite Task-Completion Dialogue Policy Learning via Hierarchical Deep Reinforcement Learning
- Resource-rational Task Decomposition to Minimize Planning Costs
- Hierarchical Reinforcement Learning for Quadruped Locomotion
- Decentralized Reinforcement Learning: Global Decision-Making via Local Economic Transactions
- TTR-Based Reward for Reinforcement Learning with Implicit Model Priors
- Deep Imitation Learning for Bimanual Robotic Manipulation
- TACO: Learning Task Decomposition via Temporal Alignment for Control
- A survey on intrinsic motivation in reinforcement learning
- The Value Equivalence Principle for Model-Based Reinforcement Learning
- Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?
- Learning and Planning in Average-Reward Markov Decision Processes
- Experience Replay with Likelihood-free Importance Weights
- Compositional planning in Markov decision processes: Temporal abstraction meets generalized logic composition
- Deep Reinforcement Learning at the Edge of the Statistical Precipice
- Exploiting Hierarchy for Learning and Transfer in KL-regularized RL
- Unifying task specification in reinforcement learning
- Disentangled Planning and Control in Vision Based Robotics via Reward Machines
- Hierarchical Decision Making In Electricity Grid Management
- SDRL: Interpretable and Data-efficient Deep Reinforcement Learning Leveraging Symbolic Planning
- Learning and Planning for Time-Varying MDPs Using Maximum Likelihood Estimation
- HAVEN: Hierarchical Cooperative Multi-Agent Reinforcement Learning with Dual Coordination Mechanism
- Substitute Teacher Networks: Learning with Almost No Supervision
- Crowdfunding Dynamics Tracking: A Reinforcement Learning Approach
- Efficient Learning and Planning with Compressed Predictive States
- Reinforcement Learning Applications
- Sub-Goal Trees -- a Framework for Goal-Directed Trajectory Prediction and Optimization
- A Boolean Task Algebra for Reinforcement Learning
- Relational Deep Reinforcement Learning for Routing in Wireless Networks
- Perception-Prediction-Reaction Agents for Deep Reinforcement Learning
- Combating the Compounding-Error Problem with a Multi-step Model
- A representational framework for learning and encoding structurally enriched trajectories in complex agent environments
- Expanding LLM Agent Boundaries with Strategy-Guided Exploration
- Learning Shared Representations in Multi-task Reinforcement Learning
- Reinforcement Learning with Competitive Ensembles of Information-Constrained Primitives
- A Spiking Neural Learning Classifier System
- Intermittent Active Inference
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation
- Bottom-Up Skill Discovery from Unsegmented Demonstrations for Long-Horizon Robot Manipulation
- Inter-Agent Relative Representations for Multi-Agent Option Discovery
- Enhancing Hierarchical Reinforcement Learning through Change Point Detection in Time Series
- Learning to Drive Safely with Hybrid Options
- Learning Parameterized Skills from Demonstrations
- Toward Agents That Reason About Their Computation
- HRM-Agent: Training a recurrent reasoning model in dynamic environments using reinforcement learning
- The Eigenoption-Critic Framework
- Sequentially Teaching Sequential Tasks (ST)2: Teaching Robots Long-horizon Manipulation Skills
- Reinforcement Learning and Consumption-Savings Behavior
- Using Large Language Models for Abstraction of Planning Domains - Extended Version
- LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments
- A Principle of Targeted Intervention for Multi-Agent Reinforcement Learning
- Logic-based Task Representation and Reward Shaping in Multiagent Reinforcement Learning
- RLAF: Reinforcement Learning from Automaton Feedback
- Expressive Reward Synthesis with the Runtime Monitoring Language
- Learn to Change the World: Multi-level Reinforcement Learning with Model-Changing Actions
- Demystifying the Mechanisms Behind Emergent Exploration in Goal-conditioned RL
- Optimistic Reinforcement Learning-Based Skill Insertions for Task and Motion Planning
- Unsupervised Learning of Object Keypoints for Perception and Control
- AI Agents as Universal Task Solvers
- Disentangling Options with Hellinger Distance Regularizer
- Homomorphic Mappings for Value-Preserving State Aggregation in Markov Decision Processes
- Decision-Theoretic Planning with Concurrent Temporally Extended Actions
- Compositional meta-learning through probabilistic task inference
- Distributed Planning in Hierarchical Factored MDPs
- A Micro-Objective Perspective of Reinforcement Learning
- Learning Transferable Concepts in Deep Reinforcement Learning
- DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
- Intelligent AI Delegation
- Motion Planner Augmented Reinforcement Learning for Robot Manipulation in Obstructed Environments
- From Pixels to Factors: Learning Independently Controllable State Variables for Reinforcement Learning
- Focused Skill Discovery: Learning to Control Specific State Variables while Minimizing Side Effects
- Learning to Reason as Action Abstractions with Scalable Mid-Training RL
- EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance
- Serotonin predictively encodes value
- Playing Atari Ball Games with Hierarchical Reinforcement Learning
- State2vec: Off-Policy Successor Features Approximators
- Inter-Level Cooperation in Hierarchical Reinforcement Learning
- Reinforcement Learning with Discrete Diffusion Policies for Combinatorial Action Spaces
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- Composing trajectories for rapid inference of navigational goals
- Autonomous learning and chaining of motor primitives using the Free Energy Principle
- ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation
- Policy Compatible Skill Incremental Learning via Lazy Learning Interface
- Guided Policy Search for Parameterized Skills using Adverbs
- Example-Driven Model-Based Reinforcement Learning for Solving Long-Horizon Visuomotor Tasks
- Hierarchical Decision Making by Generating and Following Natural Language Instructions
- AC-Teach: A Bayesian Actor-Critic Method for Policy Learning with an Ensemble of Suboptimal Teachers
- Situationally Aware Options
- Object Manipulation Learning by Imitation
- Multi-layer Abstraction for Nested Generation of Options (MANGO) in Hierarchical Reinforcement Learning
- Efficient Exploration in Constrained Environments with Goal-Oriented Reference Path
- A Deep Hierarchical Approach to Lifelong Learning in Minecraft
- ARE: Scaling Up Agent Environments and Evaluations
- When should agents explore?
- MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale
- LodeStar: Long-horizon Dexterity via Synthetic Data Augmentation from Human Demonstrations
- Symmetry Learning for Function Approximation in Reinforcement Learning
- Robust Reinforcement Learning for Continuous Control with Model Misspecification
- A Deep Value-network Based Approach for Multi-Driver Order Dispatching
- From Machine Learning to Robotics: Challenges and Opportunities for\n Embodied Intelligence
- Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
- Hierarchical Variational Imitation Learning of Control Programs
- Solving Compositional Reinforcement Learning Problems via Task Reduction
- Average-Reward Learning and Planning with Options
- Monte Carlo tree search with spectral expansion for planning with dynamical systems
- An Off-policy Policy Gradient Theorem Using Emphatic Weightings
- Action and Perception as Divergence Minimization
- HumanoidVerse: A Versatile Humanoid for Vision-Language Guided Multi-Object Rearrangement
- Hybrid Reward Architecture for Reinforcement Learning
- NurseSchedRL: Attention-Guided Reinforcement Learning for Nurse-Patient Assignment
- Goal Kernel Planning: Linearly-Solvable Non-Markovian Policies for Logical Tasks with Goal-Conditioned Options
- Physics-informed Value Learner for Offline Goal-Conditioned Reinforcement Learning
- Reinforcement Learning with Anticipation: A Hierarchical Approach for Long-Horizon Tasks
- Generative World Models of Tasks: LLM-Driven Hierarchical Scaffolding for Embodied Agents
- Self-Consistent Trajectory Autoencoder: Hierarchical Reinforcement Learning with Trajectory Embeddings
- Multi-robot task allocation through vacancy chain scheduling
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- Planning with Reasoning using Vision Language World Model
- Generative Temporal Difference Learning for Infinite-Horizon Prediction
- Off-Policy Actor-Critic with Shared Experience Replay
- Transfer Learning in Deep Reinforcement Learning: A Survey
- Hierarchical Reinforcement Learning By Discovering Intrinsic Options
- TAAC: Temporally Abstract Actor-Critic for Continuous Control
- A Framework for Constrained and Adaptive Behavior-Based Agents
- ADAPS: Autonomous Driving Via Principled Simulations
- Visualizing Dynamics: from t-SNE to SEMI-MDPs
- Neuro-Symbolic Predictive Process Monitoring
- Scalable Option Learning in High-Throughput Environments
- Few-Shot Neuro-Symbolic Imitation Learning for Long-Horizon Planning and Acting
- Interactive Agent Modeling by Learning to Probe
- An Analysis of Frame-skipping in Reinforcement Learning
- Hierarchical Reinforcement Learning for Deep Goal Reasoning: An Expressiveness Analysis
- Learning Meta Representations for Agents in Multi-Agent Reinforcement Learning
- AI Research Considerations for Human Existential Safety (ARCHES)
- HEAS: Hierarchical Evolutionary Agent-Based Simulation Framework for Multi-Objective Policy Search
- Real-Time Model Checking for Closed-Loop Robot Reactive Planning
- Benchmarking Deep Reinforcement Learning for Continuous Control
- Multitask Soft Option Learning
- Optimal Control of Complex Systems through Variational Inference with a Discrete Event Decision Process
- Routing Networks and the Challenges of Modular and Compositional Computation
- Adaptive Skip Intervals: Temporal Abstraction for Recurrent Dynamical Models
- Speeding Up Planning in Markov Decision Processes via Automatically Constructed Abstractions
- Discovering hierarchies using Imitation Learning from hierarchy aware policies
- Sub-policy Adaptation for Hierarchical Reinforcement Learning
- Landmark-Assisted Monte Carlo Planning
- Self-organization of action hierarchy and compositionality by reinforcement learning with recurrent neural networks
- Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse Reward
- Reinforcement learning in the brain
- Representational efficiency outweighs action efficiency in human program induction
- Subgoal Search For Complex Reasoning Tasks
- From proprioception to long-horizon planning in novel environments: A hierarchical RL model
- The Efficiency of Human Cognition Reflects Planned Information Processing
- Lifelong Learning using Eigentasks: Task Separation, Skill Acquisition, and Selective Transfer
- Adversarial Skill Chaining for Long-Horizon Robot Manipulation via Terminal State Regularization
- Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving
- Hierarchy through Composition with Linearly Solvable Markov Decision Processes
- Reinforcement Learning for Target Zone Blood Glucose Control
- Learning Temporal Abstractions via Variational Homomorphisms in Option-Induced Abstract MDPs
- Efficient Solving of Large Single Input Superstate Decomposable Markovian Decision Process
- Hierarchical Message-Passing Policies for Multi-Agent Reinforcement Learning
- CompILE: Compositional Imitation Learning and Execution
- Dynamics-Aware Unsupervised Discovery of Skills
- Some Insights into Lifelong Reinforcement Learning Systems
- Near-Optimal Representation Learning for Hierarchical Reinforcement Learning
- Hierarchical Reinforcement Learning with Hindsight
- Generating Adjacency-Constrained Subgoals in Hierarchical Reinforcement Learning
- Planning for Decentralized Control of Multiple Robots Under Uncertainty
- Learning Generalizable Locomotion Skills with Hierarchical Reinforcement Learning
- Hierarchical Skills for Efficient Exploration
- A Greedy Approach to Adapting the Trace Parameter for Temporal Difference Learning
- Directed Policy Gradient for Safe Reinforcement Learning with Human Advice
- Multimodal Hierarchical Reinforcement Learning Policy for Task-Oriented Visual Dialog
- TOMA: Topological Map Abstraction for Reinforcement Learning
- HARLF: Hierarchical Reinforcement Learning and Lightweight LLM-Driven Sentiment Integration for Financial Portfolio Optimization
- Towards Microgrid Resilience Enhancement via Mobile Power Sources and Repair Crews: A Multi-Agent Reinforcement Learning Approach
- Uniform State Abstraction For Reinforcement Learning
- An Empirical Comparison of Off-policy Prediction Learning Algorithms in the Four Rooms Environment
- A Matrix Splitting Perspective on Planning with Options
- A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning
- Option Discovery in Hierarchical Reinforcement Learning using Spatio-Temporal Clustering
- Modular Multitask Reinforcement Learning with Policy Sketches
- Deep Residual Reinforcement Learning
- Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous Systems
- Behavior Priors for Efficient Reinforcement Learning
- Neuro-Inspired Inverse Learning for Planning and Control
- Temporally Abstract Partial Models
- Model-based Reinforcement Learning for Semi-Markov Decision Processes with Neural ODEs
- Hierarchical Planning with Latent World Models
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
- Grounding Hierarchical Reinforcement Learning Models for Knowledge Transfer
- HILONet: Hierarchical Imitation Learning from Non-Aligned Observations
- Reinforce Lifelong Interaction Value of User-Author Pairs for Large-Scale Recommendation Systems
- Emergent Real-World Robotic Skills via Unsupervised Off-Policy Reinforcement Learning
- On the Sample Complexity of End-to-end Training vs. Semantic Abstraction Training
- SkillBlender: Towards Versatile Humanoid Whole-Body Loco-Manipulation via Skill Blending
- The Design and Implementation of XiaoIce, an Empathetic Social Chatbot
- Age of Information: An Introduction and Survey
- Learning Task Decomposition with Ordered Memory Policy Network
- Scalable and Cost-Efficient de Novo Template-Based Molecular Generation
- Hierarchical Imitation and Reinforcement Learning
- An Open-Source Framework for Adaptive Traffic Signal Control
- Learning to Locomote with Deep Neural-Network and CPG-based Control in a Soft Snake Robot
- Hierarchical Policies for Cluttered-Scene Grasping with Latent Plans
- Parrot: Data-Driven Behavioral Priors for Reinforcement Learning
- Back to Basics: Deep Reinforcement Learning in Traffic Signal Control
- Adaptable Agent Populations via a Generative Model of Policies
- Mixed Discrete and Continuous Planning using Shortest Walks in Graphs of Convex Sets
- Universal Memory Architectures for Autonomous Machines
- ELLA: Exploration through Learned Language Abstraction
- Option Discovery in the Absence of Rewards with Manifold Analysis
- Learning more skills through optimistic exploration
- Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited Data
- Learning Purposeful Behaviour in the Absence of Rewards
- Learning The Minimum Action Distance
- TRAIL: Near-Optimal Imitation Learning with Suboptimal Data
- SPLASH! Sample-efficient Preference-based inverse reinforcement learning for Long-horizon Adversarial tasks from Suboptimal Hierarchical demonstrations
- Time-Aware Q-Networks: Resolving Temporal Irregularity for Deep Reinforcement Learning
- Reinforcement Learning with Action Chunking
- Hierarchical Reinforcement Learning via Advantage-Weighted Information Maximization
- Situational Awareness by Risk-Conscious Skills
- First Return, Entropy-Eliciting Explore
- Learning to Compose Skills
- Memorize or generalize? Searching for a compositional RNN in a haystack
- Transferring Agent Behaviors from Videos via Motion GANs
- Hierarchical Reinforcement Learning with Targeted Causal Interventions
- Automatic Representation for Lifetime Value Recommender Systems
- Affordance as general value function: A computational model
- Synthesized Policies for Transfer and Adaptation across Tasks and Environments
- Active Screening for Recurrent Diseases: A Reinforcement Learning Approach
- Zero-Incentive Dynamics: a look at reward sparsity through the lens of unrewarded subgoals
- Zero-Shot Task Generalization with Multi-Task Deep Reinforcement Learning
- Learning Goal Embeddings via Self-Play for Hierarchical Reinforcement Learning
- Divide-and-Conquer Monte Carlo Tree Search For Goal-Directed Planning
- Active Hierarchical Imitation and Reinforcement Learning
- Dynamics-aware Embeddings
- Forgetful Experience Replay in Hierarchical Reinforcement Learning from Demonstrations
- Reinforcement Learning with Physics-Informed Symbolic Program Priors for Zero-Shot Wireless Indoor Navigation
- A Hierarchical Framework for Relation Extraction with Reinforcement Learning
- Advances and applications in inverse reinforcement learning: a comprehensive review
- Hierarchical reinforcement learning for efficient exploration and transfer
- REALab: An Embedded Perspective on Tampering
- Graph-Assisted Stitching for Offline Hierarchical Reinforcement Learning
- Bi-level Off-policy Reinforcement Learning for Volt/VAR Control Involving Continuous and Discrete Devices
- Piecewise-constant Neural ODEs
- Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning
- A Berkeley View of Systems Challenges for AI
- FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement Learning
- Efficient Strategy Synthesis for MDPs via Hierarchical Block Decomposition
- When Can Model-Free Reinforcement Learning be Enough for Thinking?
- Goal-conditioned Hierarchical Reinforcement Learning for Sample-efficient and Safe Autonomous Driving at Intersections
- Neural Arithmetic Expression Calculator
- Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments
- HiLight: A Hierarchical Reinforcement Learning Framework with Global Adversarial Guidance for Large-Scale Traffic Signal Control
- Efficient Exploration through Intrinsic Motivation Learning for Unsupervised Subgoal Discovery in Model-Free Hierarchical Reinforcement Learning
- Separation of Concerns in Reinforcement Learning
- A free energy principle for a particular physics
- Leveraging Sequentiality in Reinforcement Learning from a Single Demonstration
- Causality in the human niche: lessons for machine learning
- Fast Monte Carlo Tree Diffusion: 100x Speedup via Parallel Sparse Planning
- Eigenoption Discovery through the Deep Successor Representation
- Delphos: A reinforcement learning framework for assisting discrete choice model specification
- Temporally-Extended ε-Greedy Exploration
- Diversity-Driven Extensible Hierarchical Reinforcement Learning
- WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
- Hierarchical model-based policy optimization: from actions to action sequences and back
- Skill Transfer in Deep Reinforcement Learning under Morphological Heterogeneity
- Hierarchical Deep Multiagent Reinforcement Learning with Temporal Abstraction
- Avoiding Tampering Incentives in Deep RL via Decoupled Approval
- An Empirical Study of Representation Learning for Reinforcement Learning in Healthcare
- Interactive Learning of Environment Dynamics for Sequential Tasks
- Model-Based AI planning and Execution Systems for Robotics
- Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning
- Zero-Shot Generalization using Intrinsically Motivated Compositional Emergent Protocols
- Provable Hierarchy-Based Meta-Reinforcement Learning
- A Joint Planning and Learning Framework for Human-Aided Decision-Making
- How to Avoid Being Eaten by a Grue: Structured Exploration Strategies for Textual Worlds
- Ray Interference: a Source of Plateaus in Deep Reinforcement Learning
- Catastrophic Importance of Catastrophic Forgetting
- Implicitly Aligning Humans and Autonomous Agents through Shared Task Abstractions
- Meta Learning Shared Hierarchies
- Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback
- PPO in the Fisher-Rao geometry
- POMDPs in Continuous Time and Discrete Spaces
- Horizon Reduction Makes RL Scalable
- Reducing Commitment to Tasks with Off-Policy Hierarchical Reinforcement Learning
- Single-step Options for Adversary Driving
- Active Learning of Inverse Models with Intrinsically Motivated Goal Exploration in Robots
- Learning State Abstractions for Transfer in Continuous Control
- Learning agile soccer skills for a bipedal robot with deep reinforcement learning
- Return-based Scaling: Yet Another Normalisation Trick for Deep RL
- Constructing Abstraction Hierarchies Using a Skill-Symbol Loop
- Mobi-π: Mobilizing Your Robot Learning Policy
- Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents
- Hyperbolic Embeddings for Learning Options in Hierarchical Reinforcement Learning
- Deep Reinforcement Learning for Dexterous Manipulation with Concept Networks
- Exploiting Language Instructions for Interpretable and Compositional Reinforcement Learning
- Reward Propagation Using Graph Convolutional Networks
- An empirical investigation of the challenges of real-world reinforcement learning
- Hierarchical visuomotor control of humanoids
- Learning Multi-Agent Coordination for Enhancing Target Coverage in Directional Sensor Networks
- MRS: Multi-Resolution Skills for HRL Agents
- Constraint Satisfaction Propagation: Non-stationary Policy Synthesis for Temporal Logic Planning
- On a Formal Model of Safe and Scalable Self-driving Cars
- Near Optimal Behavior via Approximate State Abstraction
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
- Importance Resampling for Off-policy Prediction
- A Laplacian Framework for Option Discovery in Reinforcement Learning
- Planning in Hierarchical Reinforcement Learning: Guarantees for Using Local Policies
- Learning Reusable Options for Multi-Task Reinforcement Learning
- Knowledge Retrieval in LLM Gaming: A Shift from Entity-Centric to Goal-Oriented Graphs
- Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
- MAP Propagation Algorithm: Faster Learning with a Team of Reinforcement Learning Agents
- Physically Embedded Planning Problems: New Challenges for Reinforcement Learning
- Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit
- Integrating Counterfactual Simulations with Language Models for Explaining Multi-Agent Behaviour
- Learning Plannable Representations with Causal InfoGAN
- Predicting Periodicity with Temporal Difference Learning
- ShIOEnv: A Command Evaluation Environment for Grammar-Constrained Synthesis and Execution Behavior Modeling
- Optimal Options for Multi-Task Reinforcement Learning Under Time Constraints
- RLOC: Neurobiologically Inspired Hierarchical Reinforcement Learning Algorithm for Continuous Control of Nonlinear Dynamical Systems
- CM3: Cooperative Multi-goal Multi-stage Multi-agent Reinforcement Learning
- StableMimic: Smooth Human-Like Recovery for Humanoid Motion Tracking - Learning Beyond the Tracking Distribution for Structured Post-Fall Behavior
- Learning Graph Structure With A Finite-State Automaton Layer
- Forethought and Hindsight in Credit Assignment
- Improving planning and MBRL with temporally-extended actions
- TempoRL: Learning When to Act
- Bridging the Gap: Providing Post-Hoc Symbolic Explanations for Sequential Decision-Making Problems with Inscrutable Representations
- Reward Shaping with Subgoals for Social Navigation
- Timely CPU Scheduling for Computation-intensive Status Updates
- Flattening Hierarchies with Policy Bootstrapping
- VeRecycle: Reclaiming Guarantees from Probabilistic Certificates for Stochastic Dynamical Systems after Change
- Training Small LLMs as Spatial Multi-Agent Policies
- Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement Learning
- InnateCoder: Learning Programmatic Options with Foundation Models
- Multi-timescale Nexting in a Reinforcement Learning Robot
- Hindsight policy gradients
- Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning
- PoE-World: Compositional World Modeling with Products of Programmatic Experts
- Decoupling Dynamics and Reward for Transfer Learning
- Learning Virtual Machine Scheduling in Cloud Computing through Language Agents
- Compositional Planning Using Optimal Option Models
- Electric Bus Charging Schedules Relying on Real Data-Driven Targets Based on Hierarchical Deep Reinforcement Learning
- Goal-seeking compresses neural codes for space in the human hippocampus and orbitofrontal cortex
- DSADF: Thinking Fast and Slow for Decision Making
- Intrinsically Motivated Acquisition of Modular Slow Features for Humanoids in Continuous and Non-Stationary Environments
- Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review
- A temporally abstracted Viterbi algorithm
- Let Humanoids Hike! Integrative Skill Development on Complex Trails
- Inverse Reinforcement Learning via Nonparametric Spatio-Temporal Subgoal Modeling
- Deep Reinforcement Learning and Transportation Research: A Comprehensive Review
- Discovering Generalizable Skills via Automated Generation of Diverse Tasks
- Complex Skill Acquisition Through Simple Skill Imitation Learning
- D3HRL: A Distributed Hierarchical Reinforcement Learning Approach Based on Causal Discovery and Spurious Correlation Detection
- IV-Posterior: Inverse Value Estimation for Interpretable Policy Certificates
- Classifying Options for Deep Reinforcement Learning
- The Logical Options Framework
- Spatial Learning and Action Planning in a Prefrontal Cortical Network Model
- Hierarchical Cross-Modal Agent for Robotics Vision-and-Language Navigation
- Adjust Planning Strategies to Accommodate Reinforcement Learning Agents
- Deep Reinforcement Learning: An Overview
- Reinforcement Learning Under Algorithmic Triage
- Model Predictive Control For Trade Execution
- Learning Parameterized Skills
- Evidence of an Emergent "Self" in Continual Robot Learning
- What is intrinsic motivation? A typology of computational approaches
- Learning Hierarchical Teaching Policies for Cooperative Agents
- HAC Explore: Accelerating Exploration with Hierarchical Reinforcement Learning
- The Price Is Not Right: Neuro-Symbolic Methods Outperform VLAs on Structured Long-Horizon Manipulation Tasks with Significantly Lower Energy Consumption
- An Empirical Comparison of Off-policy Prediction Learning Algorithms on the Collision Task
- Continual Learning of Control Primitives: Skill Discovery via Reset-Games
- RL Post-Training Builds Compositional Reasoning Strategies
- Off-Policy Adversarial Inverse Reinforcement Learning
- Quantum Reinforcement Learning
- Discovery of Options via Meta-Learned Subgoals
- Bayesian Relational Memory for Semantic Visual Navigation
- Playing Doom with SLAM-Augmented Deep Reinforcement Learning
- PODNet: A Neural Network for Discovery of Plannable Options
- Multi-Objective Molecule Generation using Interpretable Substructures
- Constructing an Optimal Behavior Basis for the Option Keyboard
- Discovering Options for Exploration by Minimizing Cover Time
- Reachable Space Characterization of Markov Decision Processes with Time Variability
- Hierarchical Prompt-Domain Control and Learning for Resource-Constrained Agentic Language Models
- Cognitive maps are generative programs
- BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems
- Hierarchical Reinforcement Learning in Multi-Goal Spatial Navigation with Autonomous Mobile Robots
- Coupled Control, Structured Memory, and Verifiable Action in Agentic AI (SCRAT -- Stochastic Control with Retrieval and Auditable Trajectories): A Comparative Perspective from Squirrel Locomotion and Scatter-Hoarding
- Learning Compositional Neural Programs for Continuous Control
- Unsupervised Hierarchical Skill Discovery
- Generalized Inverse Planning: Learning Lifted non-Markovian Utility for Generalizable Task Representation
- Learning to Sequence Robot Behaviors for Visual Navigation
- DynoPlan: Combining Motion Planning and Deep Neural Network based Controllers for Safe HRL
- Neurosymbolic Reasoning with Incremental Knowledge for Sample Efficient Hierarchical Reinforcement Learning
- oIRL: Robust Adversarial Inverse Reinforcement Learning with Temporally Extended Actions
- On the Complexity of Exploration in Goal-Driven Navigation
- Accurately and Efficiently Interpreting Human-Robot Instructions of Varying Granularities
- MOSAIC: Skill-Centric Manipulation Planning with Physics Simulation
- On mechanisms for transfer using landmark value functions in multi-task\n lifelong reinforcement learning
- Solving Sokoban using Hierarchical Reinforcement Learning with Landmarks
- A Classification View on Meta Learning Bandits
- Decentralized Cooperative Planning for Automated Vehicles with Hierarchical Monte Carlo Tree Search
- Preprint: Exploring Inevitable Waypoints for Unsolvability Explanation in Hybrid Planning Problems
- Dynamic Contrastive Skill Learning with State-Transition Based Skill Clustering and Dynamic Length Adjustment
- Scheduled Intrinsic Drive: A Hierarchical Take on Intrinsically Motivated Exploration
- Dexterous Manipulation through Imitation Learning: A Survey
- Temporal-adaptive Hierarchical Reinforcement Learning
- Discrete Sequential Prediction of Continuous Actions for Deep RL
- Online Baum-Welch algorithm for Hierarchical Imitation Learning
- Approximate Exploration through State Abstraction
- Learning Hierarchical Integration of Foveal and Peripheral Vision for Vergence Control by Active Efficient Coding
- Induction and Exploitation of Subgoal Automata for Reinforcement Learning
- An Extensible Interactive Interface for Agent Design
- Online Off-policy Prediction
- State-Continuity Approximation of Markov Decision Processes via Finite Element Methods for Autonomous System Planning
- Hierarchical Latent Prediction for Language Models
- Subgoal-based Reward Shaping to Improve Efficiency in Reinforcement Learning
- TW-CRL: Time-Weighted Contrastive Reward Learning for Efficient Inverse Reinforcement Learning
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- What is Intrinsic Motivation? A Typology of Computational Approaches. [europepmc]
- Credit assignment in multiple goal embodied visuomotor behavior. [europepmc]
- Why we should talk about option generation in decision-making research. [europepmc]
- Which is the best intrinsic motivation signal for learning multiple skills? [europepmc]
- Optimal indolence: a normative microscopic approach to work and leisure. [europepmc]
- The algorithmic anatomy of model-based evaluation. [europepmc]
- Model-based hierarchical reinforcement learning and human action control. [europepmc]
- Divide et impera: subgoaling reduces the complexity of probabilistic inference and problem solving. [europepmc]
- Problem Solving as Probabilistic Inference with Subgoaling: Explaining Human Successes and Pitfalls in the Tower of Hanoi. [europepmc]
- A Unified Theoretical Framework for Cognitive Sequencing. [europepmc]
- A neural model of hierarchical reinforcement learning. [europepmc]
- Planning and navigation as active inference. [europepmc]
- Compositional clustering in task structure learning. [europepmc]
- Rational metareasoning and the plasticity of cognitive control. [europepmc]
- Hierarchical motor control in mammals and machines. [europepmc]
- Discovery of hierarchical representations for efficient planning. [europepmc]
- Generalizing to generalize: Humans flexibly switch between compositional and conjunctive structures during reinforcement learning. [europepmc]
- Sentience and the Origins of Consciousness: From Cartesian Duality to Markovian Monism. [europepmc]
- Continuous decisions. [europepmc]
- Monkey plays Pac-Man with compositional strategies and hierarchical decision-making. [europepmc]
- Humans can navigate complex graph structures acquired during latent learning. [europepmc]
- Arithmetic value representation for hierarchical behavior composition. [europepmc]
- Anxiety as a disorder of uncertainty: implications for understanding maladaptive anxiety, anxious avoidance, and exposure therapy. [europepmc]
- A Survey on Deep Reinforcement Learning Algorithms for Robotic Manipulation. [europepmc]
- Humans decompose tasks by trading off utility and computational cost. [europepmc]