2020/11/02 by Jonah Siekmann, Siekmann, Jonah, Yesh Godse +5 · 29 citations
Computer Science · Engineering · #FOS: Computer and information sciences #Muscle activation and electromyography studies #Reinforcement Learning in Robotics #Robotic Locomotion and Control #Robotics (cs.RO) #cs.RO
paper · pdf · doi:10.48550/arxiv.2011.01387
Accepted for presentation at ICRA 2021. The first two authors contributed equally to this work
openalex publication_date 2020/11/02 · arxiv created 2021/03/11 · arxiv updated 2021/03/12 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
We study the problem of realizing the full spectrum of bipedal locomotion on a real robot with sim-to-real reinforcement learning (RL). A key challenge of learning legged locomotion is describing different gaits, via reward functions, in a way that is intuitive for the designer and specific enough to reliably learn the gait across different initial random seeds or hyperparameters. A common approach is to use reference motions (e.g. trajectories of joint positions) to guide learning. However, finding high-quality reference motions can be difficult and the trajectories themselves narrowly constrain the space of learned motion. At the other extreme, reference-free reward functions are often underspecified (e.g. move forward) leading to massive variance in policy behavior, or are the product of significant reward-shaping via trial-and-error, making them exclusive to specific gaits. In this work, we propose a reward-specification framework based on composing simple probabilistic periodic costs on basic forces and velocities. We instantiate this framework to define a parametric reward function with intuitive settings for all common bipedal gaits - standing, walking, hopping, running, and skipping. Using this function we demonstrate successful sim-to-real transfer of the learned gaits to the bipedal robot Cassie, as well as a generic policy that can transition between all of the two-beat gaits.