A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
2010/11/02 by Stéphane Ross, Stephane Ross, Geoffrey J. Gordon +4 · 2 voices · 724 citations
Computer Science · Decision Sciences · Mathematics · #Advanced Bandit Algorithms Research #Machine Learning and Algorithms #Reinforcement Learning in Robotics #cs.AI #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1011.0686
Appearing in the 14th International Conference on Artificial Intelligence and Statistics (AISTATS 2011)
arxiv published 2010/11/02 · arxiv created 2011/03/16 · arxiv updated 2011/03/16
Abstract
Sequential prediction problems such as imitation learning, where future observations depend on previous predictions (actions), violate the common i.i.d. assumptions made in statistical learning. This leads to poor performance in theory and often in practice. Some recent approaches provide stronger guarantees in this setting, but remain somewhat unsatisfactory as they train either non-stationary or stochastic policies and require a large number of iterations. In this paper, we propose a new iterative algorithm, which trains a stationary deterministic policy, that can be seen as a no regret algorithm in an online learning setting. We show that any such no regret algorithm, combined with additional reduction assumptions, must find a policy with good performance under the distribution of observations it induces in such sequential settings. We demonstrate that this new approach outperforms previous approaches on two challenging imitation learning problems and a benchmark sequence labeling problem.
Cited by
- Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control
- Act2Goal: From World Model To General Goal-conditioned Policy
- A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
- MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation
- Understanding Human-like Solutions in Combinatorial Optimization via Learning and Search
- FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models
- Learned Interventions in Lean 4 grind
- CAST: Game Solvers as Turn-Level Teachers for LLM Agents
- Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features
- The Cartesian Cut in Agentic AI
- AstraNav-Memory: Contexts Compression for Long Memory
- ETP-R1: Evolving Topological Planning with Reinforcement Fine-tuning for Vision-Language Navigation in Continuous Environments
- EGM: Efficiently Learning General Motion Tracking Policy for High Dynamic Humanoid Whole-Body Control
- SD2AIL: Adversarial Imitation Learning from Synthetic Demonstrations via Diffusion Models
- Systematic Benchmarking of SUMO Against Data-Driven Traffic Simulators
- Distributionally Robust Imitation Learning: Layered Control Architecture for Certifiable Autonomy
- TakeAD: Preference-based Post-optimization for End-to-end Autonomous Driving with Expert Takeover Data
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- Simulation-Driven Railway Delay Prediction: An Imitation Learning Approach
- Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections
- History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation
- OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
- Sample-Efficient Robot Skill Learning for Construction Tasks: Benchmarking Hierarchical Reinforcement Learning and Vision-Language-Action VLA Model
- OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models
- State-Conditional Adversarial Learning: An Off-Policy Visual Domain Transfer Method for End-to-End Imitation Learning
- A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach
- User-Feedback-Driven Adaptation for Vision-and-Language Navigation
- Efficient-VLN: A Training-Efficient Vision-Language Navigation Model
- Closing the Train-Test Gap in World Models for Gradient-Based Planning
- Learning Agile Striker Skills for Humanoid Soccer Robots from Noisy Sensory Input
- Hybrid-Diffusion Models: Combining Open-loop Routines with Visuomotor Diffusion Policies
- Rethinking Camera Choice: An Empirical Study on Fisheye Camera Properties in Robotic Manipulation
- Comparison of neural network training strategies for the simulation of dynamical systems
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- ATHENA: Agentic Team for Hierarchical Evolutionary Numerical Algorithms
- GrOMP: Grasped Object Manifold Projection for Multimodal Imitation Learning of Manipulation
- Nav-R2 Dual-Relation Reasoning for Generalizable Open-Vocabulary Object-Goal Navigation
- Synthetic Error Injection Fails to Elicit Self-Correction In Language Models
- Learning Dexterous Manipulation Skills from Imperfect Simulations
- RoaD: Rollouts as Demonstrations for Closed-Loop Supervised Fine-Tuning of Autonomous Driving Policies
- Much Ado About Noising: Dispelling the Myths of Generative Robotic Control
- Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
- Sample-Efficient Expert Query Control in Active Imitation Learning via Conformal Prediction
- Beyond Egocentric Limits: Multi-View Depth-Based Learning for Robust Quadrupedal Locomotion
- U Net LSTM with incremental time-stepping for robust long-horizon unsteady flow prediction
- Consistent inverse optimal control for discrete-time nonlinear stochastic systems
- Hyper-GoalNet: Goal-Conditioned Manipulation Policy Learning with HyperNetworks
- Escaping the Verifier: Learning to Reason via Demonstrations
- Dataset Poisoning Attacks on Behavioral Cloning Policies
- Arcadia: Toward a Full-Lifecycle Framework for Embodied Lifelong Learning
- Learning Massively Multitask World Models for Continuous Control
- Neural surrogates for designing gravitational wave detectors
- SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control
- UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model
- AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
- Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation
- VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation
- π*0.6: a VLA That Learns From Experience
- FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration
- Residual Policy Learning
- Toward the Fundamental Limits of Imitation Learning
- A Minimalist Approach to Offline Reinforcement Learning
- Dialogue Learning with Human Teaching and Feedback in End-to-End Trainable Task-Oriented Dialogue Systems
- Learning to Search Better Than Your Teacher
- Leveraging Human Guidance for Deep Reinforcement Learning Tasks
- Teachable Reinforcement Learning via Advice Distillation
- Reinforcement Learning with Feedback Graphs
- FoldPath: End-to-End Object-Centric Motion Generation via Modulated Implicit Paths
- On the Entropy Calibration of Language Models
- Intuitive Programming, Adaptive Task Planning, and Dynamic Role Allocation in Human-Robot Collaboration
- The Path Not Taken: RLVR Provably Learns Off the Principals
- Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments
- Cross-domain Imitation from Observations
- Model-Predictive Policy Learning with Uncertainty Regularization for Driving in Dense Traffic
- Energy-Based Imitation Learning
- Balance Equation-based Distributionally Robust Offline Imitation Learning
- Unified Humanoid Fall-Safety Policy from a Few Demonstrations
- Action-Based Representation Learning for Autonomous Driving
- Towards Adaptive Humanoid Control via Multi-Behavior Distillation and Reinforced Fine-Tuning
- Adversarially Regularized Policy Learning Guided by Trajectory Optimization
- Imitation Learning in the Deep Learning Era: A Novel Taxonomy and Recent Advances
- Dexterous Robotic Piano Playing at Scale
- Human-in-the-loop Online Rejection Sampling for Robotic Manipulation
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
- Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models
- When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning
- Sidekick Policy Learning for Active Visual Exploration
- Imitation Learning from Pixel-Level Demonstrations by HashReward
- Imitation with Neural Density Models
- Avoidance of Manual Labeling in Robotic Autonomous Navigation Through Multi-Sensory Semi-Supervised Learning
- On the Sample Complexity of Stability Constrained Imitation Learning
- Neural Text Generation with Unlikelihood Training
- Feedback in Imitation Learning: The Three Regimes of Covariate Shift
- Reinforcement and Imitation Learning via Interactive No-Regret Learning
- Bridging the Imitation Gap by Adaptive Insubordination
- Projection-Based Constrained Policy Optimization
- Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
- Fast and Efficient Locomotion via Learned Gait Transitions
- On Value Discrepancy of Imitation Learning
- SILG: The Multi-environment Symbolic Interactive Language Grounding Benchmark
- Object-Aware Regularization for Addressing Causal Confusion in Imitation Learning
- Domain-Robust Visual Imitation Learning with Mutual Information Constraints
- Learning spatial hearing via innate mechanisms
- Sequence Level Training with Recurrent Neural Networks
- To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning
- Visual Adversarial Imitation Learning using Variational Models
- Robust Asymmetric Learning in POMDPs
- Offline Learning from Demonstrations and Unlabeled Experience
- Discriminator-Actor-Critic: Addressing Sample Inefficiency and Reward Bias in Adversarial Imitation Learning
- CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
- Learning End-to-end Multimodal Sensor Policies for Autonomous Navigation
- Curriculum Offline Imitation Learning
- Reward Models are Metrics in a Trench Coat
- MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
- One-Shot Visual Imitation Learning via Meta-Learning
- Perceptual Attention-based Predictive Control
- Computational Morphology with Neural Network Approaches
- Path Integral Networks: End-to-End Differentiable Optimal Control
- Backplay: "Man muss immer umkehren"
- Imitation Learning for Neural Morphological String Transduction
- Sub-Goal Trees -- a Framework for Goal-Directed Trajectory Prediction and Optimization
- TrafficSim: Learning to Simulate Realistic Multi-Agent Behaviors
- Social Attention for Autonomous Decision-Making in Dense Traffic
- Strictly Batch Imitation Learning by Energy-based Distribution Matching
- Combating the Compounding-Error Problem with a Multi-step Model
- Sequence-to-Sequence Learning as Beam-Search Optimization
- AutoPhase: Compiler Phase-Ordering for High Level Synthesis with Deep Reinforcement Learning
- Human and Multi-Agent collaboration in a human-MARL teaming framework
- Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving
- Automatic Curricula via Expert Demonstrations
- Learning Reductions that Really Work
- Evaluating model-based planning and planner amortization for continuous control
- f-GAIL: Learning f-Divergence for Generative Adversarial Imitation Learning
- Meta-Adversarial Inverse Reinforcement Learning for Decision-making Tasks
- Integrative neurocybernetic modeling in the era of large-scale neuroscience
- Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning?
- Fighting Copycat Agents in Behavioral Cloning from Observation Histories
- Buzz, Choose, Forget: A Meta-Bandit Framework for Bee-Like Decision Making
- NavQ: Learning a Q-Model for Foresighted Vision-and-Language Navigation
- Towards Robust and Adaptive Motion Forecasting: A Causal Representation Perspective
- Fill in the Blanks: Accelerating Q-Learning with a Handful of Demonstrations in Sparse Reward Settings
- Blindfolded Experts Generalize Better: Insights from Robotic Manipulation and Videogames
- Reinforcement Learning from Imperfect Demonstrations
- Learning Parameterized Skills from Demonstrations
- Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation
- MPC-Inspired Neural Network Policies for Sequential Decision Making
- Preventing Posterior Collapse with Levenshtein Variational Autoencoder
- Is Temporal Difference Learning the Gold Standard for Stitching in RL?
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- SutureBot: A Precision Framework & Benchmark For Autonomous End-to-End Suturing
- Dino-Diffusion Modular Designs Bridge the Cross-Domain Gap in Autonomous Parking
- GSWorld: Closed-Loop Photo-Realistic Simulation Suite for Robotic Manipulation
- Approximate Model Predictive Control for Microgrid Energy Management via Imitation Learning
- Learning to Play No-Press Diplomacy with Best Response Policy Iteration
- Perfect Prediction or Plenty of Proposals? What Matters Most in Planning for Autonomous Driving
- Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
- SPOT: Scalable Policy Optimization with Trees for Markov Decision Processes
- Data Efficient Reinforcement Learning for Legged Robots
- Embodied Navigation with Auxiliary Task of Action Description Prediction
- MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation
- R2BC: Multi-Agent Imitation Learning from Single-Agent Demonstrations
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
- RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation
- Graph Attention-Guided Search for Dense Multi-Agent Pathfinding
- Consistent Zero-Shot Imitation with Contrastive Goal Inference
- Decentralized Real-Time Planning for Multi-UAV Cooperative Manipulation via Imitation Learning
- Learning to Answer from Correct Demonstrations
- Benchmarking Model-Based Reinforcement Learning
- When Planners Meet Reality: How Learned, Reactive Traffic Agents Shift nuPlan Benchmarks
- From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
- pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
- A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control
- Learning Social Navigation from Positive and Negative Demonstrations and Rule-Based Specifications
- Diffusion Models for Reinforcement Learning: Foundations, Taxonomy, and Development
- DemoHLM: From One Demonstration to Generalizable Humanoid Loco-Manipulation
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- MLE-guided parameter search for task loss minimization in neural sequence modeling
- Deep Learning of Robotic Tasks without a Simulator using Strong and Weak Human Supervision
- Parallel Scheduled Sampling
- Interactive-Predictive Neural Machine Translation through Reinforcement and Imitation
- Failure Prediction at Runtime for Generative Robot Policies
- Toward Autonomous Soft Robotic Endovascular Navigation via Imitation Learning
- Scalable Offline Metrics for Autonomous Driving
- Agent Learning via Early Experience
- Neurosymbolic Reinforcement Learning with Formally Verified Exploration
- Learning Sparse Rewarded Tasks from Sub-Optimal Demonstrations
- Bayesian Decision Making around Experts
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- DecompGAIL: Learning Realistic Traffic Behaviors with Decomposed Multi-Agent Generative Adversarial Imitation Learning
- Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality
- Imitation Learning from Imperfect Demonstration
- On the Guaranteed Almost Equivalence between Imitation Learning from Observation and Demonstration
- Predictive Preference Learning from Human Interventions
- Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
- From Learning to Mastery: Achieving Safe and Efficient Real-World Autonomous Driving with Human-In-The-Loop Reinforcement Learning
- Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning
- Pathology-CoT: Learning Visual Chain-of-Thought Agent from Expert Whole Slide Image Diagnosis Behavior
- Challenges of Real-World Reinforcement Learning
- Neural Map: Structured Memory for Deep Reinforcement Learning
- Reward Augmented Maximum Likelihood for Neural Structured Prediction
- SITCOM: Scaling Inference-Time COMpute for VLAs
- RAP: 3D Rasterization Augmented End-to-End Planning
- Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
- A Control-Barrier-Function-Based Algorithm for Policy Adaptation in Reinforcement Learning
- Geometric Properties of Neural Multivariate Regression
- Rethinking Thinking Tokens: LLMs as Improvement Operators
- TubeDAgger: Reducing the Number of Expert Interventions with Stochastic Reach-Tubes
- Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation
- Data-Efficient Multitask DAgger
- Quantifying Generalisation in Imitation Learning
- Fidelity-Aware Data Composition for Robust Robot Generalization
- EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance
- In-Hand Manipulation of Articulated Tools with Dexterous Robot Hands with Sim-to-Real Transfer
- FTACT: Force Torque aware Action Chunking Transformer for Pick-and-Reorient Bottle Task
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- Agile perceptive multiskill locomotion for quadrupedal robots in the wild
- Learning hierarchical behavior and motion planning for autonomous driving
- Data Driven Aircraft Trajectory Prediction with Deep Imitation Learning
- JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
- Towards Target-Driven Visual Navigation in Indoor Scenes via Generative Imitation Learning
- RobotDancing: Residual-Action Reinforcement Learning Enables Robust Long-Horizon Humanoid Motion Tracking
- Selective Progress-Aware Querying for Human-in-the-Loop Reinforcement Learning
- Robot Trajectron V2: A Probabilistic Shared Control Framework for Navigation
- Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data Augmentation
- Competitive Multi-agent Inverse Reinforcement Learning with Sub-optimal Demonstrations
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- UniCross: Unified Cross-Skill Dexterous Manipulation Synthesis
- QuantWAMs: Calibrating at the Right Granularity for World Action Models
- LM-GRASP: Instance-Specific Language Models for Combinatorial Construction via Online Imitation Learning
- It's Not Just More Demos: Counterfactual Action Sensitivity Coverage for Data-Efficient Robust Robot Imitation
- Neurosymbolic Imitation Learning with Human Guidance: A Privileged Information Approach
- Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense
- DriverGym: Democratising Reinforcement Learning for Autonomous Driving
- Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation
- Affordance-based Reinforcement Learning for Urban Driving
- AC-Teach: A Bayesian Actor-Critic Method for Policy Learning with an Ensemble of Suboptimal Teachers
- Of Moments and Matching: A Game-Theoretic Framework for Closing the Imitation Gap
- Learning to Search via Retrospective Imitation
- Learning from Suboptimal Demonstration via Self-Supervised Reward Regression
- Learning to Navigate Sidewalks in Outdoor Environments
- Generalization Guarantees for Imitation Learning
- ROPA: Synthetic Robot Pose Generation for RGB-D Bimanual Data Augmentation
- Auditing Robot Learning for Safety and Compliance during Deployment
- SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
- Tensor-based Cooperative Control for Large Scale Multi-intersection Traffic Signal Using Deep Reinforcement Learning and Imitation Learning
- Regressing Word and Sentence Embeddings for Regularization of Neural Machine Translation
- Orchestrate, Generate, Reflect: A VLM-Based Multi-Agent Collaboration Framework for Automated Driving Policy Learning
- End2Race: Efficient End-to-End Imitation Learning for Real-Time F1Tenth Racing
- MCTS-EP: Empowering Embodied Planning with Online Preference Optimization
- KungfuBot2: Learning Versatile Motion Skills for Humanoid Whole-Body Control
- Learning from Observation: A Survey of Recent Advances
- Uncertainty-Based Smooth Policy Regularisation for Reinforcement Learning with Few Demonstrations
- Reinforcement Learning Agent for a 2D Shooter Game
- Demonstration-Guided Reinforcement Learning with Learned Skills
- Online Learning of Deceptive Policies under Intermittent Observation
- MIMIC-D: Multi-modal Imitation for MultI-agent Coordination with Decentralized Diffusion Policies
- SHaRe-RL: Structured, Interactive Reinforcement Learning for Contact-Rich Industrial Assembly Tasks
- Track Any Motions under Any Disturbances
- One-Shot Imitation Learning
- Behavior Foundation Model for Humanoid Robots
- TreeIRL: Safe Urban Driving with Tree Search and Inverse Reinforcement Learning
- VisuoSpatial Foresight for Multi-Step, Multi-Task Fabric Manipulation
- Hierarchical Variational Imitation Learning of Control Programs
- From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
- ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation
- Deceptive Risk Minimization: Out-of-Distribution Generalization by Deceiving Distribution Shift Detectors
- Provably Breaking the Quadratic Error Compounding Barrier in Imitation Learning, Optimally
- Robust Navigation for Racing Drones based on Imitation Learning and Modularization
- RAPTOR: A Foundation Policy for Quadrotor Control
- Synthetic vs. Real Training Data for Visual Navigation
- TrajBooster: Boosting Humanoid Whole-Body Manipulation via Trajectory-Centric Learning
- RSL-RL: A Learning Library for Robotics Research
- Robust Multi-Modal Policies for Industrial Assembly via Reinforcement Learning and Demonstrations: A Large-Scale Study
- HumanoidVerse: A Versatile Humanoid for Vision-Language Guided Multi-Object Rearrangement
- Self-Augmented Robot Trajectory: Efficient Imitation Learning via Safe Self-augmentation with Demonstrator-annotated Precision
- How well can LLMs provide planning feedback in grounded environments?
- Mitigating Covariate Shift in Imitation Learning via Offline Data Without Great Coverage
- Symmetry-Guided Multi-Agent Inverse Reinforcement Learning
- PLANS: Robust Program Learning from Neurally Inferred Specifications
- RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction
- SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards
- Learning to Synthesize Programs as Interpretable and Generalizable Policies
- Deep Reactive Policy: Learning Reactive Manipulator Motion Planning for Dynamic Environments
- Safety Meets Speed: Accelerated Neural MPC with Safety Guarantees and No Retraining
- Zero-shot Imitation Learning from Demonstrations for Legged Robot Visual Navigation
- Provable Representation Learning for Imitation Learning via Bi-level Optimization
- On the Utility of Model Learning in HRI
- Learning Multi-Stage Tasks with One Demonstration via Self-Replay
- Sequence-Level Knowledge Distillation
- End-to-End Offline Goal-Oriented Dialog Policy Learning via Policy Gradient
- Distributed Link Sparsification for Scalable Scheduling Using Graph Neural Networks (Journal Version)
- Imitation Learning Based on Disentangled Representation Learning of Behavioral Characteristics
- Dota 2 with Large Scale Deep Reinforcement Learning
- Riemannian Motion Policy Fusion through Learnable Lyapunov Function Reshaping
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- Imitation Learning: Progress, Taxonomies and Challenges
- Critic Regularized Regression
- FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback
- Model-based Deep Reinforcement Learning for Dynamic Portfolio Optimization
- DUViN: Diffusion-Based Underwater Visual Navigation via Knowledge-Transferred Depth Features
- Robotic Imitation of Human Assembly Skills Using Hybrid Trajectory and Force Learning
- Transfer Learning in Deep Reinforcement Learning: A Survey
- Robots of the Lost Arc: Self-Supervised Learning to Dynamically Manipulate Fixed-Endpoint Cables
- State Alignment-based Imitation Learning
- Sim-Real Joint Reinforcement Transfer for 3D Indoor Navigation
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- DropoutDAgger: A Bayesian Approach to Safe Imitation Learning
- ADAPS: Autonomous Driving Via Principled Simulations
- Reinforcement learning in robotics: A survey
- SECANT: Self-Expert Cloning for Zero-Shot Generalization of Visual Policies
- Learning Latent Plans from Play
- SABR: A Stable Adaptive Bitrate Framework Using Behavior Cloning Pretraining and Reinforcement Learning Fine-Tuning
- Evaluation Function Approximation for Scrabble
- Spiking Decision Transformers: Local Plasticity, Phase-Coding, and Dendritic Routing for Low-Power Sequence Control
- Guess-and-Learn (G&L): Measuring the Cumulative Error Cost of Cold-Start Adaptation
- Interactive Agent Modeling by Learning to Probe
- Verifiable Reinforcement Learning via Policy Extraction
- Causal Navigation by Continuous-time Neural Networks
- Distilling Policy Distillation
- Exploiting Policy Idling for Dexterous Manipulation
- Active Imitation Learning from Multiple Non-Deterministic Teachers: Formulation, Challenges, and Algorithms
- Information Templates: A New Paradigm for Intelligent Active Feature Acquisition
- Search-Based Credit Assignment for Offline Preference-Based Reinforcement Learning
- Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation
- Vision-Based Autonomous Car Racing Using Deep Imitative Reinforcement Learning
- Pessimism About Unknown Unknowns Inspires Conservatism
- Offline Imitation Learning upon Arbitrary Demonstrations by Pre-Training Dynamics Representations
- Hybrid Reinforcement Learning with Expert State Sequences
- Reinforcement and Imitation Learning for Diverse Visuomotor Skills
- Ecole: A Library for Learning Inside MILP Solvers
- Discovering hierarchies using Imitation Learning from hierarchy aware policies
- Predictor-Corrector Policy Optimization
- Belief-Conditioned One-Step Diffusion: Real-Time Trajectory Planning with Just-Enough Sensing
- No More Blind Spots: Learning Vision-Based Omnidirectional Bipedal Locomotion for Challenging Terrain
- Designing Interpretable Approximations to Deep Reinforcement Learning
- Addressing reward bias in Adversarial Imitation Learning with neutral reward functions
- Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse Reward
- KDPE: A Kernel Density Estimation Strategy for Diffusion Policy Trajectory Selection
- GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation
- DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation
- Data-to-text Generation by Splicing Together Nearest Neighbors
- Towards Affordance-Aware Robotic Dexterous Grasping with Human-like Priors
- Topo-boundary: A Benchmark Dataset on Topological Road-boundary Detection Using Aerial Images for Autonomous Driving
- Whole-Body Coordination for Dynamic Object Grasping with Legged Manipulators
- From Imitation to Optimization: A Comparative Study of Offline Learning for Autonomous Driving
- Agile Autonomous Driving using End-to-End Deep Imitation Learning
- Real-Time Adversarial Attacks
- Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution
- ASkDAgger: Active Skill-level Data Aggregation for Interactive Imitation Learning
- Analyzing the Impact of Multimodal Perception on Sample Complexity and Optimization Landscapes in Imitation Learning
- Learning to Multi-Task Learn for Better Neural Machine Translation
- Robustifying Learning-Augmented Caching Efficiently without Compromising 1-Consistency
- Learning Policies for Contextual Submodular Prediction
- Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation
- Safety-Aware Imitation Learning via MPC-Guided Disturbance Injection
- Aerobatic maneuvers in insect-scale flapping-wing aerial robots via deep-learned robust tube model predictive control
- No Need for Interactions: Robust Model-Based Imitation Learning using Neural ODE
- Winning Isn't Everything: Enhancing Game Development with Intelligent Agents
- Learning 6DoF Grasping Using Reward-Consistent Demonstration
- Bench2ADVLM: A Closed-Loop Benchmark for Vision-language Models in Autonomous Driving
- VPN: Visual Prompt Navigation
- An Auto-tuning Framework for Autonomous Vehicles
- Beyond Simulation: Benchmarking World Models for Planning and Causality in Autonomous Driving
- Learning Neural Parsers with Deterministic Differentiable Imitation Learning
- Learn by Observation: Imitation Learning for Drone Patrolling from Videos of A Human Navigator
- Video Generators are Robot Policies
- EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
- Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents
- Learning Temporal Strategic Relationships using Generative Adversarial Imitation Learning
- SoftDICE for Imitation Learning: Rethinking Off-policy Distribution Matching
- Retrieve-Augmented Generation for Speeding up Diffusion Policy without Additional Training
- Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation
- ThriftyDAgger: Budget-Aware Novelty and Risk Gating for Interactive Imitation Learning
- Tracking the Race Between Deep Reinforcement Learning and Imitation Learning -- Extended Version
- Imitation-Projected Programmatic Reinforcement Learning
- MLC-Agent: Cognitive Model based on Memory-Learning Collaboration in LLM Empowered Agent Simulation Environment
- Towards Mixed Optimization for Reinforcement Learning with Program Synthesis
- Navigation by Imitation in a Pedestrian-Rich Environment
- Causal Confusion in Imitation Learning
- AutoPhase: Juggling HLS Phase Orderings in Random Forests with Deep Reinforcement Learning
- Few-shot transfer of tool-use skills using human demonstrations with proximity and tactile sensing
- Recurrent Predictive State Policy Networks
- Test-time Offline Reinforcement Learning on Goal-related Experience
- Difference of Convex Functions Programming Applied to Control with Expert Data
- Training with Exploration Improves a Greedy Stack-LSTM Parser
- Learning to Play by Imitating Humans
- A Contraction Approach to Model-based Reinforcement Learning
- Scalable Perception-Action-Communication Loops with Convolutional and Graph Neural Networks
- Episodic Self-Imitation Learning with Hindsight
- The Role of Feedback Alignment in Self-Distillation
- What can AI do for me: Evaluating Machine Learning Interpretations in Cooperative Play
- Provably Efficient Imitation Learning from Observation Alone
- AEGIS: A Backup Reflex for Physical AI
- X-Nav: Learning End-to-End Cross-Embodiment Navigation for Mobile Robots
- Explaining Fast Improvement in Online Imitation Learning
- Semantic Visual Navigation by Watching YouTube Videos
- PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost
- Robust Maximum Entropy Behavior Cloning
- Universal Value Density Estimation for Imitation Learning and Goal-Conditioned Reinforcement Learning
- Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning
- Scalable Multi-Agent Inverse Reinforcement Learning via Actor-Attention-Critic
- Fighting Failures with FIRE: Failure Identification to Reduce Expert Burden in Intervention-Based Learning
- Gradient-free Policy Architecture Search and Adaptation
- Efficient Black-box Assessment of Autonomous Vehicle Safety
- Heterogeneous Robot Teams for Informative Sampling
- On Multi-Agent Learning in Team Sports Games
- Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
- Parser for Abstract Meaning Representation using Learning to Search
- Learning from Imperfect Demonstrations from Agents with Varying Dynamics
- Neural Co-state Regulator: A Data-Driven Paradigm for Real-time Optimal Control with Input Constraints
- Feedback Motion Planning for Liquid Transfer using Supervised Learning
- Long-Term Mobile Traffic Forecasting Using Deep Spatio-Temporal Neural Networks
- Hindsight Generative Adversarial Imitation Learning
- InfoGAIL: Interpretable Imitation Learning from Visual Demonstrations
- Learning Belief Representations for Imitation Learning in POMDPs
- Hierarchical Imitation and Reinforcement Learning
- Targeted Data Acquisition for Evolving Negotiation Agents
- Stability Conditions for Online Learnability
- End-to-End Training of Deep Visuomotor Policies
- Reinforcement Learning-based Visual Navigation with Information-Theoretic Regularization
- Virtual-Taobao: Virtualizing Real-world Online Retail Environment for Reinforcement Learning
- Uncertainty-Aware Constraint Learning for Adaptive Safe Motion Planning from Demonstrations
- Towards Learning to Imitate from a Single Video Demonstration
- Hierarchical Policies for Cluttered-Scene Grasping with Latent Plans
- Adversarial Imitation Learning from Incomplete Demonstrations
- Learning Calibratable Policies using Programmatic Style-Consistency
- Sufficiently Accurate Model Learning
- Evolutionary Stochastic Policy Distillation
- Bayesian Nonparametric Feature and Policy Learning for Decision-Making
- SAFARI: Safe and Active Robot Imitation Learning with Imagination
- Active Imitation Learning via Reduction to I.I.D. Active Learning
- Object-Centric Mobile Manipulation through SAM2-Guided Perception and Imitation Learning
- From Semantic Web and MAS to Agentic AI: A Unified Narrative of the Web of Agents
- Evolving Graphical Planner: Contextual Global Planning for Vision-and-Language Navigation
- From Reasoning to Super-Intelligence: A Search-Theoretic Perspective
- Human-in-the-Loop Imitation Learning using Remote Teleoperation
- Grasping with Chopsticks: Combating Covariate Shift in Model-free Imitation Learning for Fine Manipulation
- Am I Building a White Box Agent or Interpreting a Black Box Agent?
- Behavioral Exploration: Learning to Explore via In-Context Adaptation
- TRAIL: Near-Optimal Imitation Learning with Suboptimal Data
- Robot-Gated Interactive Imitation Learning with Adaptive Intervention Mechanism
- Imitation Learning via Simultaneous Optimization of Policies and Auxiliary Trajectories
- DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search
- UniTracker: Learning Universal Whole-Body Motion Tracker for Humanoid Robots
- Value from Observations: Towards Large-Scale Imitation Learning via Self-Improvement
- Concurrent Training Improves the Performance of Behavioral Cloning from Observation
- State-only Imitation with Transition Dynamics Mismatch
- Predictive-State Decoders: Encoding the Future into Recurrent Networks
- Modeling Preemptive Behaviors for Uncommon Hazardous Situations From Demonstrations
- Learning Policies for Markov Decision Processes from Data
- CueLearner: Bootstrapping and local policy adaptation from relative feedback
- Neural Autonomous Navigation with Riemannian Motion Policy
- StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
- HERO: History-Enriched Rollout Training for Long-Horizon Autoregressive Neural Operators
- Evaluating the Robustness of Collaborative Agents
- Improving Learning from Demonstrations by Learning from Experience
- Sample Efficient Imitation Learning via Reward Function Trained in Advance
- OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding
- Recall Traces: Backtracking Models for Efficient Reinforcement Learning
- Learning Algorithms for Regenerative Stopping Problems with Applications to Shipping Consolidation in Logistics
- Imitation Learning by Reinforcement Learning
- Learning by Playing - Solving Sparse Reward Tasks from Scratch
- Online Learning with Continuous Variations: Dynamic Regret and Reductions
- DexWrist: A Robotic Wrist for Constrained and Dynamic Manipulation
- Leveraging Genetic Algorithms for Efficient Demonstration Generation in Real-World Reinforcement Learning Environments
- Neuro-Symbolic Constraint Programming for Structured Prediction
- Learning Online from Corrective Feedback: A Meta-Algorithm for Robotics
- Towards End-to-End Deep Learning for Autonomous Racing: On Data Collection and a Unified Architecture for Steering and Throttle Prediction
- Advancing Learnable Multi-Agent Pathfinding Solvers with Active Fine-Tuning
- NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments
- Learning Obstacle Representations for Neural Motion Planning
- Shaping Advice in Deep Multi-Agent Reinforcement Learning
- Learning to Search for Dependencies
- Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
- Active Hierarchical Imitation and Reinforcement Learning
- Chained Predictions Using Convolutional Neural Networks
- Self-Imitation Learning via Generalized Lower Bound Q-learning
- REALab: An Embedded Perspective on Tampering
- Learning Heuristic Search via Imitation
- Robust Behavior Cloning Via Global Lipschitz Regularization
- CUPID: Curating Data your Robot Loves with Influence Functions
- How to Close Sim-Real Gap? Transfer with Segmentation!
- What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
- Deep Imitation Learning for Complex Manipulation Tasks from Virtual Reality Teleoperation
- Neural Network Memory Architectures for Autonomous Robot Navigation
- BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
- Compliant Residual DAgger: Improving Real-World Contact-Rich Manipulation with Human Corrections
- CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity
- Steering Your Diffusion Policy with Latent Space Reinforcement Learning
- To Follow or not to Follow: Selective Imitation Learning from Observations
- SafeMimic: Towards Safe and Autonomous Human-to-Robot Imitation for Mobile Manipulation
- A Survey of Deep Network Solutions for Learning Control in Robotics: From Reinforcement to Imitation
- Common Benchmarks Undervalue the Generalization Power of Programmatic Policies
- GMT: General Motion Tracking for Humanoid Whole-Body Control
- ClutterDexGrasp: A Sim-to-Real System for General Dexterous Grasping in Cluttered Scenes
- LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction
- Quizbowl: The Case for Incremental Question Answering
- Human-assisted Robotic Policy Refinement via Action Preference Optimization
- Learning to Reach Goals via Iterated Supervised Learning
- From Experts to a Generalist: Toward General Whole-Body Control for Humanoid Robots
- RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
- SAIL: Faster-than-Demonstration Execution of Imitation Learning Policies
- MultiNet: Multi-Modal Multi-Task Learning for Autonomous Driving
- On Value Functions and the Agent-Environment Boundary
- Optimizing Medical Treatment for Sepsis in Intensive Care: from Reinforcement Learning to Pre-Trial Evaluation
- Behavior Planning at Urban Intersections through Hierarchical Reinforcement Learning
- Comparing Human-Centric and Robot-Centric Sampling for Robot Deep Learning from Demonstrations
- Sim-to-Real Transfer of Accurate Grasping with Eye-In-Hand Observations and Continuous Control
- Environment Reconstruction with Hidden Confounders for Reinforcement Learning based Recommendation
- Continuous Doubly Constrained Batch Reinforcement Learning
- Fixing exposure bias with imitation learning needs powerful oracles
- Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative Sampling
- Policy Learning Using Weak Supervision
- Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to Search
- Risk-Averse Offline Reinforcement Learning
- Provably Efficient Third-Person Imitation from Offline Observation
- End-to-end Learning of Driving Models from Large-scale Video Datasets
- Robust Imitation of Diverse Behaviors
- Accelerating Safe Reinforcement Learning with Constraint-mismatched Policies
- LORM: Learning to Optimize for Resource Management in Wireless Networks with Few Training Samples
- Avoiding Tampering Incentives in Deep RL via Decoupled Approval
- Relational Mimic for Visual Adversarial Imitation Learning
- DMRL: Data- and Model-aware Reward Learning for Data Extraction
- Distilling Motion Planner Augmented Policies into Visual Control Policies for Robot Manipulation
- Efficient Model-Free Reinforcement Learning Using Gaussian Process
- Learning Visual Robotic Control Efficiently with Contrastive Pre-training and Data Augmentation
- Reward function shape exploration in adversarial imitation learning: an empirical study
- Interactive Learning of Environment Dynamics for Sequential Tasks
- Self-Supervised Learning of State Estimation for Manipulating Deformable Linear Objects
- Game-theoretic Modeling of Traffic in Unsignalized Intersection Network for Autonomous Vehicle Control Verification and Validation
- Adversarial Learning of Task-Oriented Neural Dialog Models
- Distilling Knowledge for Search-based Structured Prediction
- Executing Instructions in Situated Collaborative Interactions
- iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks
- Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving
- Benders Cut Classification via Support Vector Machines for Solving Two-stage Stochastic Programs
- The Game Imitation: Deep Supervised Convolutional Networks for Quick Video Game AI
- CoT is Not True Reasoning, It Is Just a Tight Constraint to Imitate: A Theory Perspective
- Think Twice, Act Once: A Co-Evolution Framework of LLM and RL for Large-Scale Decision Making
- ACNMP: Skill Transfer and Task Extrapolation through Learning from Demonstration and Reinforcement Learning via Representation Sharing
- Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
- ADEPT: Adaptive Diffusion Environment for Policy Transfer Sim-to-Real
- HoMeR: Learning In-the-Wild Mobile Manipulation via Hybrid Imitation and Whole-Body Control
- Re-determinizing Information Set Monte Carlo Tree Search in Hanabi
- Deep Value Model Predictive Control
- Event Extraction with Generative Adversarial Imitation Learning
- Interactive Imitation Learning for Dexterous Robotic Manipulation: Challenges and Perspectives -- A Survey
- Investigating the Effects of Robot Engagement Communication on Learning from Demonstration
- AMBER: Adaptive Mesh Generation by Iterative Mesh Resolution Prediction
- LocoTouch: Learning Dynamic Quadrupedal Transport with Tactile Sensing
- Human-in-the-Loop Deep Reinforcement Learning with Application to Autonomous Driving
- Autonomy 2.0: Why is self-driving always 5 years away?
- Computational Algebra with Attention: Transformer Oracles for Border Basis Algorithms
- Regret Minimization for Partially Observable Deep Reinforcement Learning
- An empirical investigation of the challenges of real-world reinforcement learning
- Efficient Controllable Diffusion via Optimal Classifier Guidance
- Beyond Templates: Dynamic Adaptation of Reasoning Demonstrations via Feasibility-Aware Exploration
- TWIST: Teleoperated Whole-Body Imitation System
- Loss Functions for Multiset Prediction
- Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
- Gait-Conditioned Reinforcement Learning with Multi-Phase Curriculum for Humanoid Locomotion
- Deep Object-Centric Policies for Autonomous Driving
- Simultaneous Mapping and Target Driven Navigation
- Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
- SAM: Squeeze-and-Mimic Networks for Conditional Visual Driving Policy Learning
- ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving
- Inductive Certificate Synthesis for Control Design
- Inverse Q-Learning Done Right: Offline Imitation Learning in Qπ-Realizable MDPs
- SMAP: Self-supervised Motion Adaptation for Physically Plausible Humanoid Whole-body Control
- Making Teams and Influencing Agents: Efficiently Coordinating Decision Trees for Interpretable Multi-Agent Reinforcement Learning
- Designing Pin-pression Gripper and Learning its Dexterous Grasping with Online In-hand Adjustment
- MaskedManipulator: Versatile Whole-Body Manipulation
- Language Model Distillation: A Temporal Difference Imitation Learning Perspective
- Perception-and-action system for humanoid robot task execution in construction
- DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion
- CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning
- Sequential Graph Dependency Parser
- Provably Efficient Generative Adversarial Imitation Learning for Online and Offline Setting with Linear Function Approximation
- Learning from Demonstration with Weakly Supervised Disentanglement
- CoNav-UAV: Cooperative Dual-Altitude Aerial Navigation via Stackelberg Learning
- Model-based Adversarial Imitation Learning
- Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
- PLAS: Latent Action Space for Offline Reinforcement Learning
- Learning Dynamic Feature Selection for Fast Sequential Prediction
- Deep Reinforcement Learning from Policy-Dependent Human Feedback
- Playing Minecraft with Behavioural Cloning
- Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation
- Adversarial Imitation via Variational Inverse Reinforcement Learning
- How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
- Amortized Bethe Free Energy Minimization for Learning MRFs
- Efficient and Interpretable Robot Manipulation with Graph Neural Networks
- Agnostic System Identification for Model-Based Reinforcement Learning
- Guided Policy Optimization under Partial Observability
- Boosting Offline Reinforcement Learning with Residual Generative Modeling
- Convergence of Value Aggregation for Imitation Learning
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- When to retrain a machine learning model
- Transfer Learning for Mixed-Integer Resource Allocation Problems in Wireless Networks
- Latent Programmer: Discrete Latent Codes for Program Synthesis
- Embodied Passive Aeroacoustic Perception Enables Relative Sensing and Pursuit Between Aerial Robots
- Multi-granularity Textual Adversarial Attack with Behavior Cloning
- Autoregressive Knowledge Distillation through Imitation Learning
- Model Imitation for Model-Based Reinforcement Learning
- MTIL: Encoding Full History with Mamba for Temporal Imitation Learning
- Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning
- When Do Surrogate Updates Improve Decisions? A Local Theory of Trajectory-Wise Transfer
- L2D2: Robot Learning from 2D Drawings
- Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation
- Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation
- The Gate, Not the Cache: Gate Provenance Bounds the Closed-Loop Reliability of Training-Free VLA Token Skipping
- One-Shot Hierarchical Imitation Learning of Compound Visuomotor Tasks
- Unleashing Humanoid Reaching Potential via Real-world-Ready Skill Space
- Accelerating Visual-Policy Learning through Parallel Differentiable Simulation
- Hierarchical Surgical Robot Transformer (SRT-H): Imitation Learning for Autonomous Surgery
- Learning Long-Context Diffusion Policies via Past-Token Prediction
- Distilling Realizable Students from Unrealizable Teachers
- Sequential Dynamic Decision Making with Deep Neural Nets on a Test-Time Budget
- Continuous Online Learning and New Insights to Online Imitation Learning
- Preference Optimization for Combinatorial Optimization Problems
- Meta learning Framework for Automated Driving
- Self-Imitation Learning by Planning
- DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies
- HuB: Learning Extreme Humanoid Balance
- Learning to Reason and Navigate: Parameter Efficient Action Planning with Large Language Models
- Learning Robust Bed Making using Deep Imitation Learning with DART
- Learning Structural Edits via Incremental Tree Transformations
- ACORN: Adaptive Contrastive Optimization for Safe and Robust Fine-Grained Robotic Manipulation
- HyperGuide: Hyperbolic Guidance for Efficient Multi-Step Reasoning in Large Language Models
- Let Humanoids Hike! Integrative Skill Development on Complex Trails
- SOD: Step-wise On-policy Distillation for Small Language Model Agents
- CubeDAgger: Interactive Imitation Learning for Dynamic Systems with Efficient yet Low-risk Interaction
- CCL: Collaborative Curriculum Learning for Sparse-Reward Multi-Agent Reinforcement Learning via Co-evolutionary Task Evolution
- D-CODA: Diffusion for Coordinated Dual-Arm Data Augmentation
- Fast Policy Learning through Imitation and Reinforcement
- Prompt-responsive Object Retrieval with Memory-augmented Student-Teacher Learning
- Sample-Efficient Imitation Learning via Generative Adversarial Nets
- Third-Person Imitation Learning
- Smooth Imitation Learning for Online Sequence Prediction
- Imitation Learning via Off-Policy Distribution Matching
- Iterative Reinforcement Learning Based Design of Dynamic Locomotion Skills for Cassie
- Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement
- Real-Time Web Scale Event Summarization Using Sequential Decision Making
- Learning Hierarchical Teaching Policies for Cooperative Agents
- Inverse reinforcement learning for autonomous navigation via differentiable semantic mapping and planning
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- Light-Loco-Parkour: Versatile Perceptive Whole-Body Locomotion via Multi-Skill Distillation
- Text Generation by Learning from Demonstrations
- Learning to Optimize via Wasserstein Deep Inverse Optimal Control
- Safe end-to-end imitation learning for model predictive control
- Active Information Acquisition
- Adaptive Approximate Policy Iteration
- Design and Control of Roller Grasper V2 for In-Hand Manipulation
- A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
- Constrained Text Generation with Global Guidance -- Case Study on CommonGen
- Tactile Genesis: Exploring Tactile Sensors at Scale for Learning Dexterous Tasks
- PODNet: A Neural Network for Discovery of Plannable Options
- Scaling Self-Play for End-to-End Driving
- Learning Composable Energy Surrogates for PDE Order Reduction
- To Go or Not To Go? A Near Unsupervised Learning Approach For Robot Navigation
- RoboPocket: Improve Robot Policies Instantly with Your Phone
- Dynamic Oracle for Neural Machine Translation in Decoding Phase
- Design Space of Behaviour Planning for Autonomous Driving
- RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
- Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning
- Towards a Neural Debugger for Python
- Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching
- Trust-Region Behavior Blending for On-Policy Distillation
- When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
- LangWBC: Language-directed Humanoid Whole-Body Control via End-to-end Learning
- LLM-based Interactive Imitation Learning for Robotic Manipulation
- Self-Distilled Agentic Reinforcement Learning
- Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?
- OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies
- Language-Critique Imitation Learning from Suboptimal Demonstrations
- DanceOPD: On-Policy Generative Field Distillation
- When Outcome Looks Right But Discipline Fails: Trace-Based Evaluation Under Hidden Competitor State
- OmniXtreme: Breaking the Generality Barrier in High-Dynamic Humanoid Control
- Maximum Likelihood Reinforcement Learning
- χ0: Resource-Aware Robust Manipulation via Taming Distributional Inconsistencies
- Toward Fully Autonomous Driving: AI, Challenges, Opportunities, and Needs
- Fauna Sprout: A lightweight, approachable, developer-ready humanoid robot
- From Routes to Steps: Separating Semantic Progress from Local Execution in Vision-and-Language Navigation
- Agentic Reinforcement Learning with Self-Distilled Reward Shaping
- Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging
- EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning
- Tired Actor: Fatigue-Informed Character Control
- Support-weighted Adversarial Imitation Learning
- Integrating Learning-Based Manipulation and Physics-Based Locomotion for Whole-Body Badminton Robot Control
- Learning Equational Theorem Proving
- ImitAL: Learning Active Learning Strategies from Synthetic Data
- Sim2Real View Invariant Visual Servoing by Recurrent Control
- Data Scaling Laws for End-to-End Autonomous Driving
- SPECI: Skill Prompts based Hierarchical Continual Imitation Learning for Robot Manipulation
- Trajectory-based Learning for Ball-in-Maze Games
- Multimodal Perception for Goal-oriented Navigation: A Survey
- Policy Message Passing: A New Algorithm for Probabilistic Graph Inference
- SuFIA-BC: Generating High Quality Demonstration Data for Visuomotor Policy Learning in Surgical Subtasks
- When Search Teaches Style: Causal Internalization of Tactical Priors in AlphaZero
- Learning Through Retrospection: Improving Trajectory Prediction for Automated Driving with Error Feedback
- Provable Representation Learning for Imitation with Contrastive Fourier Features
- Online Baum-Welch algorithm for Hierarchical Imitation Learning
- A Model-Based Approach to Imitation Learning through Multi-Step Predictions
- Meta-Learning for Contextual Bandit Exploration
- Bayesian Learning-Based Adaptive Control for Safety Critical Systems
- An Extensible Interactive Interface for Agent Design
- Smooth Imitation Learning via Smooth Costs and Smooth Policies
- ADAPT: Actively Discovering and Adapting to Preferences for any Task
- Towards Forceful Robotic Foundation Models: a Literature Survey
- Training a Conditioned Video Game Agent on a VLM Annotated Dataset
- Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models
- Looking beyond the next token
- A Clean Slate for Offline Reinforcement Learning
- PPF: Pre-training and Preservative Fine-tuning of Humanoid Locomotion via Model-Assumption-based Regularization
- FLoRA: Sample-Efficient Preference-based RL via Low-Rank Style Adaptation of Reward Functions
- FIGARO: Hierarchical Reinforcement Learning for Scalable Microservice Management in the Computing Continuum
- Investigating the Treacherous Turn in Deep Reinforcement Learning
- Learning Goal-Oriented Visual Dialog Agents: Imitating and Surpassing Analytic Experts
- Stratified Expert Cloning for Retention-Aware Recommendation at Scale
- Model-Agnostic Policy Explanations with Large Language Models
Discussions
Related