2020/11/02 by Jonah Siekmann, Siekmann, Jonah, Yesh Godse +5 · 18 citations
Computer Science · Engineering · #FOS: Computer and information sciences #Muscle activation and electromyography studies #Reinforcement Learning in Robotics #Robotic Locomotion and Control #Robotics (cs.RO)
paper · pdf · doi:10.48550/arxiv.2011.01387
openalex publication_date 2020/11/02 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
We study the problem of realizing the full spectrum of bipedal locomotion on\na real robot with sim-to-real reinforcement learning (RL). A key challenge of\nlearning legged locomotion is describing different gaits, via reward functions,\nin a way that is intuitive for the designer and specific enough to reliably\nlearn the gait across different initial random seeds or hyperparameters. A\ncommon approach is to use reference motions (e.g. trajectories of joint\npositions) to guide learning. However, finding high-quality reference motions\ncan be difficult and the trajectories themselves narrowly constrain the space\nof learned motion. At the other extreme, reference-free reward functions are\noften underspecified (e.g. move forward) leading to massive variance in policy\nbehavior, or are the product of significant reward-shaping via trial-and-error,\nmaking them exclusive to specific gaits. In this work, we propose a\nreward-specification framework based on composing simple probabilistic periodic\ncosts on basic forces and velocities. We instantiate this framework to define a\nparametric reward function with intuitive settings for all common bipedal gaits\n- standing, walking, hopping, running, and skipping. Using this function we\ndemonstrate successful sim-to-real transfer of the learned gaits to the bipedal\nrobot Cassie, as well as a generic policy that can transition between all of\nthe two-beat gaits.\n