vix.ing · top · new · best · stats · spec

Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy Matching

2022/02/04 by Yecheng Jason Ma, Ma, Yecheng Jason, Andrew Shen +5 · 4 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Fuel Cells and Related Materials #Machine Learning (cs.LG) #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.2202.02433

openalex publication_date 2022/02/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We propose State Matching Offline DIstribution Correction Estimation (SMODICE), a novel and versatile regression-based offline imitation learning (IL) algorithm derived via state-occupancy matching. We show that the SMODICE objective admits a simple optimization procedure through an application of Fenchel duality and an analytic solution in tabular MDPs. Without requiring access to expert actions, SMODICE can be effectively applied to three offline IL settings: (i) imitation from observations (IfO), (ii) IfO with dynamics or morphologically mismatched expert, and (iii) example-based reinforcement learning, which we show can be formulated as a state-occupancy matching problem. We extensively evaluate SMODICE on both gridworld environments as well as on high-dimensional offline benchmarks. Our results demonstrate that SMODICE is effective for all three problem settings and significantly outperforms prior state-of-art.

Cited by

Related