2020/06/12 by Yunhao Tang, Tang, Yunhao · 2 citations
Computer Science · Mathematics · #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2006.07442
Accepted at NeurIPS (Neural Information Processing Systems) 2020, Vancouver, Canada. Code is available at https://github.com/robintyh1/nstep-sil
openalex publication_date 2020/06/12 · arxiv created 2021/02/14 · arxiv updated 2021/02/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Self-imitation learning motivated by lower-bound Q-learning is a novel and effective approach for off-policy learning. In this work, we propose a n-step lower bound which generalizes the original return-based lower-bound Q-learning, and introduce a new family of self-imitation learning algorithms. To provide a formal motivation for the potential performance gains provided by self-imitation learning, we show that n-step lower bound Q-learning achieves a trade-off between fixed point bias and contraction rate, drawing close connections to the popular uncorrected n-step Q-learning. We finally show that n-step lower bound Q-learning is a more robust alternative to return-based self-imitation learning and uncorrected n-step, over a wide range of continuous control benchmark tasks.