Sim-to-Real: Learning Agile Locomotion For Quadruped Robots
2018/04/27 by Jie Tan, Tan, Jie, Tingnan Zhang +14 · 114 citations
Computer Science · Engineering · #Agile software development #Artificial intelligence #Computer science #Control engineering #Engineering #Gait #Human–computer interaction #Process (computing) #Reinforcement Learning in Robotics #Reinforcement learning #Robot #Robot Manipulation and Learning #Robotic Locomotion and Control #Robotics #Simulation #Software engineering #cs.AI #cs.RO
paper · pdf · doi:10.48550/arxiv.1804.10332
published in arXiv (Cornell University) (Cornell University) · Accompanying video: https://www.youtube.com/watch?v=lUZUr7jxoqM
openalex publication_date 2018/04/27 · arxiv created 2018/05/16 · arxiv updated 2018/05/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/08
Abstract
Designing agile locomotion for quadruped robots often requires extensive expertise and tedious manual tuning. In this paper, we present a system to automate this process by leveraging deep reinforcement learning techniques. Our system can learn quadruped locomotion from scratch using simple reward signals. In addition, users can provide an open loop reference to guide the learning process when more control over the learned gait is needed. The control policies are learned in a physics simulator and then deployed on real robots. In robotics, policies trained in simulation often do not transfer to the real world. We narrow this reality gap by improving the physics simulator and learning robust policies. We improve the simulation using system identification, developing an accurate actuator model and simulating latency. We learn robust controllers by randomizing the physical environments, adding perturbations and designing a compact observation space. We evaluate our system on two agile locomotion gaits: trotting and galloping. After learning in simulation, a quadruped robot can successfully perform both gaits in the real world.
Citations
Cited by
- Egocentric Station Holding of Robotic Fish in Unknown Turbulent Background Flow
- Synthetic Data Pipelines for Adaptive, Mission-Ready Militarized Humanoids
- Learning to Get Up Across Morphologies: Zero-Shot Recovery with a Unified Humanoid Policy
- Hardware-Software Collaborative Computing of Photonic Spiking Reinforcement Learning for Robotic Continuous Control
- Beyond Egocentric Limits: Multi-View Depth-Based Learning for Robust Quadrupedal Locomotion
- Kinematics-Aware Multi-Policy Reinforcement Learning for Force-Capable Humanoid Loco-Manipulation
- VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation
- Learning Quantized Continuous Controllers for Integer Hardware
- Cross-Platform Learnable Fuzzy Gain-Scheduled Proportional-Integral-Derivative Controller Tuning via Physics-Constrained Meta-Learning and Reinforcement Learning Adaptation
- Sim-to-Real Transfer in Deep Reinforcement Learning for Bipedal Locomotion
- PHUMA: Physically-Grounded Humanoid Locomotion Dataset
- Real-DRL: Teach and Learn in Reality
- Hierarchical Reinforcement Learning for Quadruped Locomotion
- Reinforcement Learning on Cost-Constrained Quadrupedal Hardware
- Fast and Efficient Locomotion via Learned Gait Transitions
- Circus ANYmal: A Quadruped Learning Dexterous Manipulation with Its Limbs
- Sim-to-Real Transfer of Robot Learning with Variable Length Inputs
- Towards Quadrupedal Jumping and Walking for Dynamic Locomotion using Reinforcement Learning
- Toward Humanoid Brain-Body Co-design: Joint Optimization of Control and Morphology for Fall Recovery
- The Reality Gap in Robotics: Challenges, Solutions, and Best Practices
- GaussGym: An open-source real-to-sim framework for learning locomotion from pixels
- Adaptive Legged Locomotion via Online Learning for Model Predictive Control
- High-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning with Gaussian Splatting
- PolySim: Bridging the Sim-to-Real Gap for Humanoid Control via Multi-Simulator Dynamics Randomization
- DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation via Joint-Wise Neural Dynamics Model
- Sampling Strategies for Robust Universal Quadrupedal Locomotion Policies
- Learning to Act Through Contact: A Unified View of Multi-Task Robot Learning
- Agile perceptive multiskill locomotion for quadrupedal robots in the wild
- RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- qiBullet, a Bullet-based simulator for the Pepper and NAO robots
- Central Limit Theorems for Asynchronous Averaged Q-Learning
- Autonomous UAV Flight Navigation in Confined Spaces: A Reinforcement Learning Approach
- Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
- Self-Improving Embodied Foundation Models
- SoK: Cybersecurity Assessment of Humanoid Ecosystem
- Dynamic Adaptive Legged Locomotion Policy via Decoupling Reaction Force Control and Gait Control
- Reinforcement Learning with Adaptive Curriculum Dynamics Randomization for Fault-Tolerant Robot Control
- Balance Between Efficient and Effective Learning: Dense2Sparse Reward Shaping for Robot Manipulation with Environment Uncertainty
- Can SSD-Mamba2 Unlock Reinforcement Learning for End-to-End Motion Control?
- Learning agile and dynamic motor skills for legged robots
- Learning to See before Learning to Act: Visual Pre-training for Manipulation
- Transfer Learning in Deep Reinforcement Learning: A Survey
- Non-conflicting Energy Minimization in Reinforcement Learning based Robot Control
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design
- QuadKAN: KAN-Enhanced Quadruped Motion Control via End-to-End Reinforcement Learning
- Adaptive Power System Emergency Control using Deep Reinforcement Learning
- Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids
- LocoMamba: Vision-Driven Locomotion via End-to-End Deep Reinforcement Learning with Mamba
- Unsupervised Skill Discovery as Exploration for Learning Agile Locomotion
- Multimodal Attention Branch Network for Perspective-Free Sentence Generation
- Optimizing Bipedal Locomotion for The 100m Dash With Comparison to Human Running
- Learning Generalizable Locomotion Skills with Hierarchical Reinforcement Learning
- Planning in Learned Latent Action Spaces for Generalizable Legged Locomotion
- Similarity metrics for Different Market Scenarios in Abides
- Online Algorithms and Policies Using Adaptive and Machine Learning Approaches
- Behavior-Guided Actor-Critic: Improving Exploration via Learning Policy Behavior Representation for Deep Reinforcement Learning
- Iteratively Learning Muscle Memory for Legged Robots to Master Adaptive and High Precision Locomotion
- Model-free Reinforcement Learning for Model-based Control: Towards Safe, Interpretable and Sample-efficient Agents
- Self-driving scale car trained by Deep reinforcement learning
- Solving Challenging Control Problems Using Two-Staged Deep Reinforcement Learning
- PAMTRI: Pose-Aware Multi-Task Learning for Vehicle Re-Identification Using Highly Randomized Synthetic Data
- Advances, challenges, and opportunities for legged robots
- Evaluating Robots Like Human Infants: A Case Study of Learned Bipedal Locomotion
- Evaluating the Robustness of Collaborative Agents
- GEM: Group Enhanced Model for Learning Dynamical Control Systems
- SonoGym: High Performance Simulation for Challenging Surgical Tasks with Robotic Ultrasound
- Learning a Contact-Adaptive Controller for Robust, Efficient Legged Locomotion
- Real-Time Execution of Action Chunking Flow Policies
- Learning to Dock: A Simulation-based Study on Closing the Sim2Real Gap in Autonomous Underwater Docking
- Learning Accurate Whole-body Throwing with High-frequency Residual Policy and Pullback Tube Acceleration
- Differentiable Simulation of Hard Contacts with Soft Gradients for Learning and Control
- SpikePingpong: Spike Vision-based Fast-Slow Pingpong Robot System
- ReLMoGen: Leveraging Motion Generation in Reinforcement Learning for Mobile Manipulation
- Multi-Agent Manipulation via Locomotion using Hierarchical Sim2Real
- Statistical Guarantees for Offline Domain Randomization
- AI Simulation by Digital Twins: Systematic Survey, Reference Framework, and Mapping to a Standardized Architecture
- Realizing Text-Driven Motion Generation on NAO Robot: A Reinforcement Learning-Optimized Control Pipeline
- Dynamics Randomization Revisited:A Case Study for Quadrupedal Locomotion
- Bicycle Acrobatics with Reinforcement Learning
- Deep Learned Path Planning via Randomized Reward-Linked-Goals and Potential Space Applications
- DiffCoTune: Differentiable Co-Tuning for Cross-domain Robot Control
- Learning coordinated badminton skills for legged manipulators
- HACL: History-Aware Curriculum Learning for Fast Locomotion
- H2-COMPACT: Human-Humanoid Co-Manipulation via Adaptive Contact Trajectory Policies
- McARL:Morphology-Control-Aware Reinforcement Learning for Generalizable Quadrupedal Locomotion
- Bridging the Sim-to-Real Gap in Parallel-Link Leg Mechanisms via Simulator-Side Dynamics Normalization
- Sampling-Based System Identification with Active Exploration for Legged Robot Sim2Real Learning
- Learning Impact-Rich Rotational Maneuvers via Centroidal Velocity Rewards and Sim-to-Real Techniques: A One-Leg Hopper Flip Case Study
- Towards Embodiment Scaling Laws in Robot Locomotion
- Reinforcement Learning and Control of a Lower Extremity Exoskeleton for Squat Assistance
- DexDrummer: In-Hand, Contact-Rich, and Long-Horizon Dexterous Robot Drumming
- A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
- Learning Compositional Neural Programs for Continuous Control
- High-Performance Reinforcement Learning on Spot: Optimizing Simulation Parameters with Distributional Measures
- Merging Deterministic Policy Gradient Estimations with Varied Bias-Variance Tradeoff for Effective Deep Reinforcement Learning
- Coordinating Spinal and Limb Dynamics for Enhanced Sprawling Robot Mobility
- Sim-to-Real of Humanoid Locomotion Policies via Joint Torque Space Perturbation Injection
Related