2022/05/06 by Arash Vahabpour, Tianyi Wang, Vahabpour, Arash +7 · 3 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Reinforcement Learning in Robotics #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2205.03484
openalex publication_date 2022/05/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Imitation learning is the task of replicating expert policy from demonstrations, without access to a reward function. This task becomes particularly challenging when the expert exhibits a mixture of behaviors. Prior work has introduced latent variables to model variations of the expert policy. However, our experiments show that the existing works do not exhibit appropriate imitation of individual modes. To tackle this problem, we adopt an encoder-free generative model for behavior cloning (BC) to accurately distinguish and imitate different modes. Then, we integrate it with GAIL to make the learning robust towards compounding errors at unseen states. We show that our method significantly outperforms the state of the art across multiple experiments.