Planning and acting in partially observable stochastic domains
1998/05/01 by Leslie Pack Kaelbling, Michael L. Littman, Anthony R. Cassandra · 4,161 citations
Computer Science · Mathematics · #AI-based Problem Solving and Planning #Computer science #Controller (irrigation) #Formal Methods in Verification #Logic, Reasoning, and Knowledge #Markov decision process #Markov process #Mathematical optimization #Mathematics #Observable #Partially observable Markov decision process #Work (physics)
paper · pdf · doi:10.1016/s0004-3702(98)00023-x
published in Artificial Intelligence 101(1-2), 99-134 (Elsevier BV)
openalex publication_date 1998/05/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Citations
Cited by
- A Survey of Inverse Reinforcement Learning: Challenges, Methods and Progress
- Same World, Differently Given: History-Dependent Perceptual Reorganization in Artificial Agents
- TRW: TRACE-RealWorld---An Auditable Consistency Contract for World Models as Materialized Views
- Adaptive Undulatory Locomotion of Snake-like Robots in Dynamic Viscous Environments via Deep Reinforcement Learning
- Cycles of Discourse, Speech Dysfluency, and Active Inference
- Pretraining Recurrent Networks without Recurrence
- Modeling the Development of Cellular Exhaustion and Tumor-Immune Stalemate
- Dreamer-CPC: Message Learning with World Models for Decentralized Multi-agent Reinforcement Learning
- AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning
- Breaking Feedback-Blindness: Utility-Augmented Transformer for Sequential Decision Making
- A Hardware-oriented Approach for Efficient Bayesian Inference Computation and Deployment
- Generalised Bellman recurrence and three dualities in sequential decision-making
- Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
- Toward Learning POMDPs Beyond Full-Rank Actions and State Observability
- Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making
- RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
- Expected Free Energy as Belief-Dependent Utility for rho-POMDPs
- Beyond Tracking or Shortcut: Composition-Bounded Predictive States in Poker Autoregressive Models
- Learning Abstractions for Hierarchical Planning in Program-Synthesis Agents
- Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
- Breaking the Performance Ceiling in Reinforcement Learning requires Inference Strategies
- TALES: Text Adventure Learning Environment Suite
- QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
- Partially Observable Markov Decision Processes (POMDPs) and Robotics
- Self-play for Data Efficient Language Acquisition
- Meta reinforcement learning as task inference
- Reinforcement Learning via Self-Distillation
- Assisted Perception: Optimizing Observations to Communicate State
- Heterogeneity in Multi-Agent Reinforcement Learning
- Environmental benefits of enhanced surveillance technology on airport departure operations
- How Are Learned Perception-Based Controllers Impacted by the Limits of Robust Control?
- Active inference and agency: optimal control without cost functions
- Learning Optimal Decision Making for an Industrial Truck Unloading Robot using Minimal Simulator Runs
- Batch Belief Trees for Motion Planning Under Uncertainty
- Compositional Behavioral Semantics for State Abstraction in Reinforcement Learning
- Active Learning for Autonomous Intelligent Agents: Exploration, Curiosity, and Interaction
- POMP: Pomcp-based Online Motion Planning for active visual search in indoor environments
- Interpretable Model-based Hierarchical Reinforcement Learning using Inductive Logic Programming
- Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
- MEMENTO: Memory-Guided Memetic Code-as-Policy Evolution
- What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents
- PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments
- Variational Neural Belief Parameterizations for Robust Dexterous Grasping under Multimodal Uncertainty
- RoboSafe: Safeguarding Embodied Agents via Executable Safety Logic
- Generalised Linear Models in Deep Bayesian RL with Learnable Basis Functions
- Cruising the Spectrum: Joint Spectrum Mobility and Antenna Array Management for Mobile (cm/mm)Wave Connectivity
- Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
- LLM-based Few-Shot Early Rumor Detection with Imitation Agent
- Bounty Hunter: Autonomous, Comprehensive Emulation of Multi-Faceted Adversaries
- Efficient Online Learning for Optimizing Value of Information: Theory and Application to Interactive Troubleshooting
- Learning to Track Dynamic Targets in Partially Known Environments
- Efficient Planning under Partial Observability with Unnormalized Q Functions and Spectral Learning
- Dynamic Homophily with Imperfect Recall: Modeling Resilience in Adversarial Networks
- B-ActiveSEAL: Scalable Uncertainty-Aware Active Exploration with Tightly Coupled Localization-Mapping
- Learning Finite-State Controllers for Partially Observable Environments
- AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management
- A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms
- On Learning to Think: Algorithmic Information Theory for Novel\n Combinations of Reinforcement Learning Controllers and Recurrent Neural World\n Models
- Reinforcement Learning for Monetary Policy Under Macroeconomic Uncertainty: Analyzing Tabular and Function Approximation Methods
- Elephants Don’t Pack Groceries: Robot Task Planning for Low Entropy Belief States
- Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
- Quantifying Memory Use in Reinforcement Learning with Temporal Range
- Exploiting Causality for Selective Belief Filtering in Dynamic Bayesian Networks (Extended Abstract)
- Termination Analysis of Probabilistic Programs through Positivstellensatz's
- Meta-Learning Multi-armed Bandits for Beam Tracking in 5G and 6G Networks
- Orbitofrontal Cortex as a Cognitive Map of Task Space
- Risk-Aware Reasoning for Autonomous Vehicles
- POrTAL: Plan-Orchestrated Tree Assembly for Lookahead
- Inference-Based Strategy Alignment for General-Sum Differential Games
- From monoliths to modules: Decomposing transducers for efficient world modelling
- Deep Learning in Neural Networks: An Overview
- Forecasting in Offline Reinforcement Learning for Non-stationary Environments
- Real-World Reinforcement Learning of Active Perception Behaviors
- Symmetries at the origin of hierarchical emergence
- Off-policy Evaluation in Infinite-Horizon Reinforcement Learning with\n Latent Confounders
- From CAD to POMDP: Probabilistic Planning for Robotic Disassembly of End-of-Life Products
- HAVEN: Hierarchical Adversary-aware Visibility-Enabled Navigation with Cover Utilization using Deep Transformer Q-Networks
- Fault-Tolerant MARL for CAVs under Observation Perturbations for Highway On-Ramp Merging
- Automated Generation of MDPs Using Logic Programming and LLMs for Robotic Applications
- A Computable Game-Theoretic Framework for Multi-Agent Theory of Mind
- POMDP-Based Routing for DTNs with Partial Knowledge and Dependent Failures
- Learning Massively Multitask World Models for Continuous Control
- The Complexity of Decentralized Control of Markov Decision Processes
- MARL-CC: A Mathematical Framework forMulti-Agent Reinforcement Learning in ConnectedAutonomous Vehicles: Addressing Nonlinearity,Partial Observability, and Credit Assignment forOptimal Control
- Solving POMDPs by Searching the Space of Finite Policies
- SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
- Revealing POMDPs: Qualitative and Quantitative Analysis for Parity Objectives
- DualSMC: Tunneling Differentiable Filtering and Planning under Continuous POMDPs
- More Than Irrational: Modeling Belief-Biased Agents
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning
- Reinforcement Learning for Charging Optimization of Inhomogeneous Dicke Quantum Batteries
- Graph Neural Networks, Deep Reinforcement Learning and Probabilistic Topic Modeling for Strategic Multiagent Settings
- CoRL-MPPI: Enhancing MPPI With Learnable Behaviours For Efficient And Provably-Safe Multi-Robot Collision Avoidance
- UAV-Assisted Resilience in 6G and Beyond Network Energy Saving: A Multi-Agent DRL Approach
- Learning to Focus: Prioritizing Informative Histories with Structured Attention Mechanisms in Partially Observable Reinforcement Learning
- Next-Latent Prediction Transformers Learn Compact World Models
- Partially observable Markov decision processes for spoken dialog systems
- Learning Interactive World Model for Object-Centric Reinforcement Learning
- Challenges in Credit Assignment for Multi-Agent Reinforcement Learning in Open Agent Systems
- Vectorized Online POMDP Planning
- Graph-Enhanced Policy Optimization in LLM Agent Training
- Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models
- Completion ≠ Collaboration: Scaling Collaborative Effort with Agents
- Inverse Rational Control with Partially Observable Continuous Nonlinear Dynamics
- Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions
- Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
- Hidden Incentives for Auto-Induced Distributional Shift
- Efficient Inference in Markov Control Problems
- Ecological Semantics: Programming Environments for Situated Language Understanding
- A Tutorial on Sparse Gaussian Processes and Variational Inference
- Collapsing Bandits and Their Application to Public Health Interventions
- Spectrum Exploration and Exploitation for Cognitive Radio: Recent Advances
- INVIGORATE: Interactive Visual Grounding and Grasping in Clutter
- Tensor Decomposition for Multi-agent Predictive State Representation
- Online Learning and Planning in Partially Observable Domains without Prior Knowledge
- Robust Asymmetric Learning in POMDPs
- Integrating Media Selection and Media Effects Using Decision Theory
- Partially Observable Markov Decision Process Modelling for Assessing Hierarchies
- Assurance-Scoped Reliability for Agentic Networks: Capturing the State That Matters
- Learning Backward Transport for Source Localization
- A tutorial on partially observable Markov decision processes
- Efficient Learning and Planning with Compressed Predictive States
- Probabilistic Reasoning about Actions in Nonmonotonic Causal Theories
- Perception-Prediction-Reaction Agents for Deep Reinforcement Learning
- Learning to Gather Information via Imitation
- DQN-TAMER: Human-in-the-Loop Reinforcement Learning with Intractable Feedback
- Neural Co-state Policies: Structuring Hidden States in Recurrent Reinforcement Learning
- Scale-invariant temporal history (SITH): optimal slicing of the past in an uncertain world
- Scalable Planning and Learning for Multiagent POMDPs: Extended Version
- Learning Latent Representations to Influence Multi-Agent Interaction
- Shaping Belief States with Generative Environment Models for RL
- Minimal Computational Preconditions for Subjective Perspective in Artificial Agents
- An output-sensitive algorithm for multi-parametric LCPs with sufficient matrices
- Monopoly Deal: A Benchmark Environment for Bounded One-Sided Response Games
- Memory-based control with recurrent neural networks
- Multi-Task Reinforcement Learning with Context-based Representations
- Multi-Environment POMDPs: Discrete Model Uncertainty Under Partial Observability
- Deep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and Navigation
- The Formalism-Implementation Gap in Reinforcement Learning Research
- HRM-Agent: Training a recurrent reasoning model in dynamic environments using reinforcement learning
- Predictive Coding Enhances Meta-RL To Achieve Interpretable Bayes-Optimal Belief Representation Under Partial Observability
- Enhancing Tactile-based Reinforcement Learning for Robotic Control
- ESCORT: Efficient Stein-variational and Sliced Consistency-Optimized Temporal Belief Representation for POMDPs
- Reinforcement Learning and Consumption-Savings Behavior
- Many Agent Reinforcement Learning Under Partial Observability
- Reward-Free Attacks in Multi-Agent Reinforcement Learning
- RLAF: Reinforcement Learning from Automaton Feedback
- Active Measuring in Reinforcement Learning With Delayed Negative Effects
- Sparse Stochastic Finite-State Controllers for POMDPs
- GammaZero: Learning To Guide POMDP Belief Space Search With Graph Representations
- Reducing Belief Deviation in Reinforcement Learning for Active Reasoning
- Diffusion Models for Reinforcement Learning: Foundations, Taxonomy, and Development
- A Flexible Multi-Agent Deep Reinforcement Learning Framework for Dynamic Routing and Scheduling of Latency-Critical Services
- MatterDoor: Sampling Zero-shot Spatio-semantic Priors using Generative Models
- Reinforcement Learning with Action-Triggered Observations
- Multi-Modal Mutual Information (MuMMI) Training for Robust Self-Supervised Deep Reinforcement Learning
- Termination of Nondeterministic Recursive Probabilistic Programs
- Qualitative MDPs and POMDPs: An Order-Of-Magnitude Approximation
- Scene Memory Transformer for Embodied Agents in Long-Horizon Tasks
- Active Learning for UAV-based Semantic Mapping
- SoK: Measuring What Matters for Closed-Loop Security Agents
- Combining Deep Reinforcement Learning and Search for Imperfect-Information Games
- ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
- The Assistive Multi-Armed Bandit
- Neural Bayesian Filtering
- Monte Carlo Motion Planning for Robot Trajectory Optimization Under Uncertainty
- MAGIC-MASK: Multi-Agent Guided Inter-Agent Collaboration with Mask-Based Explainability for Reinforcement Learning
- Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access
- Kinodynamic Motion Planning for Mobile Robot Navigation across Inconsistent World Models
- Adversarial Reinforcement Learning Framework for ESP Cheater Simulation
- ByteSized32Refactored: Towards an Extensible Interactive Text Games Corpus for LLM World Modeling and Evaluation
- Reinforcement Learning in Partially Observable Markov Decision Processes using Hybrid Probabilistic Logic Programs
- Predicting individual learning trajectories in zebrafish via the free-energy principle
- Partially Observed Structural Causal Models
- MDP modeling for multi-stage stochastic programs
- When Can AI Models Explain Learning? Validity Criteria for AI as Cognitive Models in Education
- Toward a Computational Phenomenology of Meditative Deconstruction: “Letting Go” and the Deconstruction of Experience With Active Inference
- On Optimal Control of Discounted Cost Infinite-Horizon Markov Decision Processes Under Local State Information Structures
- DemoGrasp: Universal Dexterous Grasping from a Single Demonstration
- Model-Based Reinforcement Learning under Random Observation Delays
- Leaf it to renewal: Improved predictive maintenance policies via renewal theory and decision trees
- Learning to Lead Themselves: Agentic AI in MAS using MARL
- Learning Robust Penetration Testing Policies under Partial Observability: A systematic evaluation
- Branching out: Prognostics-Based Replacement Policies for Series Systems
- On the Importance of Multistability for Horizon Generalization in Reinforcement Learning
- Self-Modification of Policy and Utility Function in Rational Agents
- The Topological Trouble With Transformers
- Probing Dec-POMDP Reasoning in Cooperative MARL
- What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty
- Robust Reinforcement Learning in POMDPs with Incomplete and Noisy Observations
- Using Indirect Encoding of Multiple Brains to Produce Multimodal Behavior
- Prepare Before You Act: Learning From Humans to Rearrange Initial States
- I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
- GUI-ReWalk: Massive Data Generation for GUI Agent via Stochastic Exploration and Intent-Aware Reasoning
- A Formal Analysis and Taxonomy of Task Allocation in Multi-Robot Systems
- Model-Based Bayesian Reinforcement Learning in Large Structured Domains
- Robust feedback motion planning via contraction theory
- Compositional shield synthesis for safe reinforcement learning in partial observability
- TextWorld: A Learning Environment for Text-based Games
- Solving the Min-Max Multiple Traveling Salesmen Problem via Learning-Based Path Generation and Optimal Splitting
- Bayesian Action Decoder for Deep Multi-Agent Reinforcement Learning
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- Real-time reinforcement learning for turbulent state-dependent control in a bluff-body wake
- WebSight: A Vision-First Architecture for Robust Web Agents
- Information-Theoretic Methods for Planning and Learning in Partially Observable Markov Decision Processes
- Optimal Immunization Policy Using Dynamic Programming
- Instance based Generalization in Reinforcement Learning
- A coalgebraic perspective on predictive processing
- A Markov Decision Process Model for Intrusion Tolerance Problems
- Language (Re)modelling: Towards Embodied Language Understanding
- Towards bridging the gap: Systematic sim-to-real transfer for diverse legged robots
- Stabilizing Deep Q-Learning with ConvNets and Vision Transformers under Data Augmentation
- Planning under Uncertainty to Goal Distributions
- Tracking Drift-Plus-Penalty: Utility Maximization for Partially Observable and Controllable Networks
- Neural Recursive Belief States in Multi-Agent Reinforcement Learning
- Estimating Disentangled Belief about Hidden State and Hidden Task for Meta-RL
- Disentangled Multi-Context Meta-Learning: Unlocking robust and Generalized Task Learning
- Online reinforcement learning with sparse rewards through an active inference capsule
- Confounding-Robust Policy Evaluation in Infinite-Horizon Reinforcement Learning
- Recurrent Off-policy Baselines for Memory-based Continuous Control
- NiceWebRL: a Python library for human subject experiments with reinforcement learning environments
- Attention on flow control: transformer-based reinforcement learning for lift regulation in highly disturbed flows
- AS2FM: Enabling Statistical Model Checking of ROS 2 Systems for Robust Autonomy
- Language and Experience: A Computational Model of Social Learning in Complex Tasks
- Goals and the Structure of Experience
- When is Particle Filtering Efficient for Planning in Partially Observed Linear Dynamical Systems?
- Decoupling Exploration and Exploitation for Meta-Reinforcement Learning\n without Sacrifices
- Ethics-Aware Safe Reinforcement Learning for Rare-Event Risk Control in Interactive Urban Driving
- SACBP: Belief Space Planning for Continuous-Time Dynamical Systems via Stochastic Sequential Action Control
- Belief-Conditioned One-Step Diffusion: Real-Time Trajectory Planning with Just-Enough Sensing
- Counting to Explore and Generalize in Text-based Games
- Posterior Sampling for Large Scale Reinforcement Learning
- Simulating Autonomous Driving in Massive Mixed Urban Traffic
- FHHOP: A Factored Hybrid Heuristic Online Planning Algorithm for Large POMDPs
- Timeline-based planning: Expressiveness and Complexity
- Robust Remote Reinforcement Learning over Unreliable Communication Channels using Homomorphic State Encoding
- ROLL: Visual Self-Supervised Reinforcement Learning with Object Reasoning
- In-Context Reinforcement Learning via Communicative World Models
- Synthetic POMDPs to Challenge Memory-Augmented RL: Memory Demand Structure Modeling
- Exploiting Agent and Type Independence in Collaborative Graphical Bayesian Games
- Fast reinforcement learning for decentralized MAC optimization
- Probabilistic Planning by Probabilistic Programming
- Transfer Learning Across Patient Variations with Hidden Parameter Markov Decision Processes
- PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement Learning
- Recurrent Value Functions
- Inferring Reward Machines and Transition Machines from Partially Observable Markov Decision Processes
- Hierarchical Message-Passing Policies for Multi-Agent Reinforcement Learning
- A Bayesian Approach to Identifying Representational Errors
- Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains
- Learning from Humans as an I-POMDP
- Common Information based Approximate State Representations in Multi-Agent Reinforcement Learning
- Planning for Decentralized Control of Multiple Robots Under Uncertainty
- Augmenting Knowledge through Statistical, Goal-oriented Human-Robot Dialog
- Multi-Agent Belief Sharing through Autonomous Hierarchical Multi-Level Clustering
- Variational Recurrent Models for Solving Partially Observable Control Tasks
- Hybrid quantum-classical algorithm for near-optimal planning in POMDPs
- Learning the Preferences of Ignorant, Inconsistent Agents
- POMDP-lite for Robust Robot Planning under Uncertainty
- Remembering the Markov Property in Cooperative MARL
- EPSILON: An Efficient Planning System for Automated Vehicles in Highly Interactive Environments
- Myopic Bounds for Optimal Policy of POMDPs: An extension of Lovejoy's structural results
- Framing Human-Robot Task Communication as a POMDP
- Learning State-Tracking from Code Using Linear RNNs
- Concentration-Bound Analysis for Probabilistic Programs and Probabilistic Recurrence Relations
- Grounding Hierarchical Reinforcement Learning Models for Knowledge Transfer
- A Framework for Decision-Theoretic Planning I: Combining the Situation Calculus, Conditional Plans, Probability and Utility
- Reinforcement Learning for Strategic Recommendations
- Offline RL With Resource Constrained Online Deployment
- MELD: Meta-Reinforcement Learning from Images via Latent State Models
- Assessing Adaptive World Models in Machines with Novel Games
- Incentive Decision Processes
- Funnel Libraries for Real-Time Robust Feedback Motion Planning
- Shared Autonomy via Hindsight Optimization for Teleoperation and Teaming
- Partially Observable Reference Policy Programming: Solving POMDPs Sans Numerical Optimisation
- FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making
- Toward Large-Scale Agent Guidance in an Urban Taxi Service
- Decision-Theoretic Coordination and Control for Active Multi-Camera Surveillance in Uncertain, Partially Observable Environments
- Bayesian Reinforcement Learning: A Survey
- An informative path planning framework for UAV-based terrain monitoring
- Learning Belief Representations for Imitation Learning in POMDPs
- Link Prediction using Embedded Knowledge Graphs
- Forward and Backward Simulations for Partially Observable Probability
- A Geometric Perspective on Self-Supervised Policy Adaptation
- Bounded Policy Synthesis for POMDPs with Safe-Reachability Objectives
- Probabilistic Loss and its Online Characterization for Simplified Decision Making Under Uncertainty
- A Survey on Reinforcement Learning Methods in Character Animation
- Bayesian Nonparametric Feature and Policy Learning for Decision-Making
- Attention-based Learning for 3D Informative Path Planning
- Robust Active Perception via Data-association aware Belief Space\n planning
- Re4MPC: Reactive Nonlinear MPC for Multi-model Motion Planning via Deep Reinforcement Learning
- MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
- Solving POMDPs by Searching in Policy Space
- Learning Adaptive Exploration Strategies in Dynamic Environments Through Informed Policy Regularization
- Reinforcement Learning using Guided Observability
- Combining Offline Models and Online Monte-Carlo Tree Search for Planning from Scratch
- Deep Learning based Urban Vehicle Trajectory Analytics
- POPCORN: Partially Observed Prediction COnstrained ReiNforcement Learning
- A Survey of Learning in Multiagent Environments: Dealing with Non-Stationarity
- Distributed Constraint Problems for Utilitarian Agents with Privacy Concerns, Recast as POMDPs
- Memory Allocation in Resource-Constrained Reinforcement Learning
- Learning Causal State Representations of Partially Observable Environments
- A Tour of Reinforcement Learning: The View from Continuous Control
- Control Synthesis in Partially Observable Environments for Complex Perception-Related Objectives
- Belief dynamics extraction
- Monte Carlo Bayesian Reinforcement Learning
- Apprenticeship Learning for Model Parameters of Partially Observable Environments
- Single Episode Policy Transfer in Reinforcement Learning
- A Markov Decision Process Analysis of the Cold Start Problem in Bayesian Information Filtering
- Inverse Reinforcement Learning in Swarm Systems
- Deep Reinforcement Learning amidst Lifelong Non-Stationarity
- Balancing Performance and Human Autonomy with Implicit Guidance Agent
- Ego-centric Learning of Communicative World Models for Autonomous Driving
- Language-Informed Synthesis of Rational Agent Models for Grounded Theory-of-Mind Reasoning On-The-Fly
- Enhancing Text-based Reinforcement Learning Agents with Commonsense Knowledge
- Spatial Language Understanding for Object Search in Partially Observed City-scale Environments
- Stochastic Shortest Path with Energy Constraints in POMDPs
- Zero-Shot Reinforcement Learning Under Partial Observability
- AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
- Active Digital Twins via Active Inference
- Individual vs. Joint Perception: a Pragmatic Model of Pointing as Communicative Smithian Helping
- Off-Policy Evaluation in Partially Observed Markov Decision Processes under Sequential Ignorability
- Partially Observable Markov Decision Processes with Behavioral Norms
- EgoMap: Projective mapping and structured egocentric memory for Deep RL
- Hierarchical Robot Navigation in Novel Environments using Rough 2-D Maps
- Optimality and robustness in path-planning under initial uncertainty
- Understanding the Origin of Information-Seeking Exploration in Probabilistic Objectives for Control
- The Complexity of Decentralized Control of Markov Decision Processes
- Internet Congestion Control via Deep Reinforcement Learning
- GRIP: Generative Robust Inference and Perception for Semantic Robot Manipulation in Adversarial Environments
- Neuro-Symbolic Reinforcement Learning with First-Order Logic
- New Approaches for Almost-Sure Termination of Probabilistic Programs
- Distributed Estimation using Bayesian Consensus Filtering
- Multi-agent Bayesian Deep Reinforcement Learning for Microgrid Energy Management under Communication Failures
- Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to Search
- WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
- Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation
- Multi-Agent Decentralized Belief Propagation on Graphs
- An Empirical Study of Representation Learning for Reinforcement Learning in Healthcare
- Interactive Language Learning by Question Answering
- Temporal Difference Variational Auto-Encoder
- Deep reinforcement learning driven inspection and maintenance planning under incomplete information and constraints
- Identification of Unexpected Decisions in Partially Observable Monte-Carlo Planning: a Rule-Based Approach
- Non-Gaussian SLAP: Simultaneous Localization and Planning Under Non-Gaussian Uncertainty in Static and Dynamic Environments
- Graph Convolutional Memory using Topological Priors
- Policy-contingent abstraction for robust robot control
- Learning Symbolic Persistent Macro-Actions for POMDP Solving Over Time
- Minimum-Latency FEC Design with Delayed Feedback: Mathematical Modeling and Efficient Algorithms
- POMDPs in Continuous Time and Discrete Spaces
- SAVOR: Skill Affordance Learning from Visuo-Haptic Perception for Robot-Assisted Bite Acquisition
- AI-Native Brand Identity: From Visual Recognition to Cryptographic Verification
- Tru-POMDP: Task Planning Under Uncertainty via Tree of Hypotheses and Open-Ended POMDPs
- Descriptive History Representations: Learning Representations by Answering Questions
- Improving Generalization in Meta-RL with Imaginary Tasks from Latent\n Dynamics Mixture
- Data-assimilated model-informed reinforcement learning
- A Scalable Method for Solving High-Dimensional Continuous POMDPs Using Local Approximation
- Reinforcement learning approach for robustness analysis of complex networks with incomplete information
- Sparse Imagination for Efficient Visual World Model Planning
- Overcoming Multi-step Complexity in Multimodal Theory-of-Mind Reasoning: A Scalable Bayesian Planner
- Optimistic critics can empower small actors
- DYSTIL: Dynamic Strategy Induction with Large Language Models for Reinforcement Learning
- Reinforcement Learning with Temporal Logic Constraints for Partially-Observable Markov Decision Processes
- Implications of Human Irrationality for Reinforcement Learning
- Counterfactual equivalence for POMDPs, and underlying deterministic environments
- Lazy Heuristic Search for Solving POMDPs with Expensive-to-Compute Belief Transitions
- Variational Inference for Data-Efficient Model Learning in POMDPs
- A Convolution and Attention Based Encoder for Reinforcement Learning under Partial Observability
- A Dynamic Reliability-Aware Service Placement for Network Function Virtualization (NFV)
- Oryx: a Scalable Sequence Model for Many-Agent Coordination in Offline MARL
- Lifted Forward Planning in Relational Factored Markov Decision Processes with Concurrent Actions
- iCORPP: Interleaved Commonsense Reasoning and Probabilistic Planning on Robots
- A Survey of Knowledge-based Sequential Decision Making under Uncertainty
- Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
- Situationally-Aware Dynamics Learning
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- State Entropy Maximization with Random Encoders for Efficient Exploration
- Deep Active Inference Agents for Delayed and Long-Horizon Environments
- Learning and Reasoning for Robot Sequential Decision Making under Uncertainty
- A Theoretical Connection Between Statistical Physics and Reinforcement Learning
- The Cell Must Go On: Agar.io for Continual Reinforcement Learning
- ChainVLA: Chaining Vision-Language-Action Queries through a Unified Execution State for Long-Horizon Manipulation
- Bayesian Policy Optimization for Model Uncertainty
- Search in Imperfect Information Games
- What is Going on Inside Recurrent Meta Reinforcement Learning Agents?
- Enter the Void - Planning to Seek Entropy When Reward is Scarce
- Sequential Monte Carlo for Policy Optimization in Continuous POMDPs
- Guided Policy Optimization under Partial Observability
- ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
- Vid2World: Crafting Video Diffusion Models to Interactive World Models
- Control of Probabilistic Systems under Dynamic, Partially Known Environments with Temporal Logic Specifications
- Rethinking System Health Management
- The Traitors: Deception and Trust in Multi-Agent Language Model Simulations
- MTIL: Encoding Full History with Mamba for Temporal Imitation Learning
- Behavior Synthesis via Contact-Aware Fisher Information Maximization
- Hierarchical Reinforcement Learning as a Model of Human Task Interleaving
- Internal State Estimation in Groups via Active Information Gathering
- Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
- Distilling Realizable Students from Unrealizable Teachers
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
- Automatic Curriculum Learning for Driving Scenarios: Towards Robust and Efficient Reinforcement Learning
- Backward Approximate Dynamic Programming with Hidden Semi-Markov Stochastic Models in Energy Storage Optimization
- Constant-Memory Strategies in Stochastic Games: Best Responses and Equilibria
- Factorization Regret mediates compositional generalization in latent space
- Provable Distributional Value Iteration under Partial Observability
- LLM-Guided Probabilistic Program Induction for POMDP Model Estimation
- The Advantage of Cross Entropy over Entropy in Iterative Information Gathering
- Working Memory Graphs
- Structure Learning in Motor Control:A Deep Reinforcement Learning Model
- KnowRob: A knowledge processing infrastructure for cognition-enabled robots
- Representations for robot knowledge in the KnowRob framework
- Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment
- Bayesian Policy Search for Stochastic Domains
- Optimal Radio Frequency Energy Harvesting with Limited Energy Arrival Knowledge
- Adversarial observations in probabilistic State-Space Models for robust Reinforcement Learning
- Bayesian updates from coalgebraic determinisation
- Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation
- MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
- Belief at Risk: Quantifying Agentic AI Model Risk with LLM-Inferred Bayesian State Filters
- Risk-Aware Planning for Transit Desert Remediation Under Demand Uncertainty
- Hyperbolic Discounting and Learning over Multiple Horizons
- Understanding Domain Randomization for Sim-to-real Transfer
- Representations and solutions for game-theoretic problems
- Minimax-Optimal Policy Regret in Partially Observable Markov Games
- Cognitive maps are generative programs
- Adaptive AI Delegation under Uncertainty: A Bayesian Governance Policy for Sequential Decision Authority
- Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents
- When Outcome Looks Right But Discipline Fails: Trace-Based Evaluation Under Hidden Competitor State
- 3D-Belief: Embodied Belief Inference via Generative 3D World Modeling
- The Auton Agentic AI Framework
- Human-level control through deep reinforcement learning
- An Imperfect Verifier is Good Enough: Learning with Noisy Rewards
- Coupled Control, Structured Memory, and Verifiable Action in Agentic AI (SCRAT -- Stochastic Control with Retrieval and Auditable Trajectories): A Comparative Perspective from Squirrel Locomotion and Scatter-Hoarding
- Generalized Inverse Planning: Learning Lifted non-Markovian Utility for Generalizable Task Representation
- LADDER: A Human-Level Bidding Agent for Large-Scale Real-Time Online Auctions
- On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning
- POMDPs for Autonomous Science Exploration
- Learning and Reasoning for Robot Dialog and Navigation Tasks
- Learning Attentive Neural Processes for Planning with Pushing Actions
- Robust Control under Stationary Ambiguity
- Belief Space Planning for Mobile Robots with Range Sensors using iLQG
- Optimal adaptive inspection and maintenance planning for deteriorating structural systems
- Learning enables adaptation in cooperation for multi-player stochastic games
- Finding Approximate POMDP solutions Through Belief Compression
- DESPOT: Online POMDP Planning with Regularization
- Safe and Effective Picking Paths in Clutter given Discrete Distributions of Object Poses
- A model of risk and mental state shifts during social interaction
- Multimodal Perception for Goal-oriented Navigation: A Survey
- Searching for a source without gradients: how good is infotaxis and how to beat it
- Optimal Decision-Making in Mixed-Agent Partially Observable Stochastic Environments via Reinforcement Learning
- AI in a vat: Fundamental limits of efficient world modelling for agent sandboxing and interpretability
- A Framework for Objective-Driven Dynamical Stochastic Fields
- Valuating Surface Surveillance Technology for Collaborative Multiple-Spot Control of Airport Departure Operations
- MPTP: Motion-planning-aware task planning for navigation in belief space
- Hidden Markov Models and Their Application for Predicting Failure Events
- Induction and Exploitation of Subgoal Automata for Reinforcement Learning
- Constrained Active Classification Using Partially Observable Markov Decision Processes
- AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
- Quantum-Structured World Models (QSWMs) for Predictive Latent Dynamics
- Timeline of artificial intelligence [wikipedia]
- The COACH prompting system to assist older adults with dementia through handwashing: an efficacy study. [europepmc]
- Reinforcement learning: Computational theory and biological mechanisms. [europepmc]
- Decision making under uncertainty: a neural model based on partially observable markov decision processes. [europepmc]
- The development of an adaptive upper-limb stroke rehabilitation robotic system. [europepmc]
- The nature of belief-directed exploratory choice in human decision-making. [europepmc]
- A computational framework for the study of confidence in humans and animals. [europepmc]
- The algorithmic anatomy of model-based evaluation. [europepmc]
- Monte Carlo Planning Method Estimates Planning Horizons during Interactive Social Exchange. [europepmc]
- Prospective Optimization with Limited Resources. [europepmc]
- Exploratory decision-making as a function of lifelong experience, not cognitive decline. [europepmc]
- A Probability Distribution over Latent Causes, in the Orbitofrontal Cortex. [europepmc]
- Reward-based training of recurrent neural networks for cognitive and value-based tasks. [europepmc]
- Predicting explorative motor learning using decision-making and motor noise. [europepmc]
- Heuristic and optimal policy computations in the human brain during sequential decision-making. [europepmc]
- A model of risk and mental state shifts during social interaction. [europepmc]
- Rational metareasoning and the plasticity of cognitive control. [europepmc]
- Multi-step planning of eye movements in visual search. [europepmc]
- Bayesian inference with incomplete knowledge explains perceptual confidence and its deviations from accuracy. [europepmc]
- Decision prioritization and causal reasoning in decision hierarchies. [europepmc]
- Optimism and pessimism in optimised replay. [europepmc]
- Epistemic Communities under Active Inference. [europepmc]
- Alternation emerges as a multi-modal strategy for turbulent odor navigation. [europepmc]
- Putting perception into action with inverse optimal control for continuous psychophysics. [europepmc]
- Reward Maximization Through Discrete Active Inference. [europepmc]
- Emergence of belief-like representations through reinforcement learning. [europepmc]