HumanPlus: Humanoid Shadowing and Imitation from Humans
2024/06/15 by Zipeng Fu, Fu, Zipeng, Qingqing Zhao +7 · 113 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Human Pose and Action Recognition #Machine Learning (cs.LG) #Robotics (cs.RO) #Systems and Control (eess.SY) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2406.10454
openalex publication_date 2024/06/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
One of the key arguments for building robots that have similar form factors to human beings is that we can leverage the massive human data for training. Yet, doing so has remained challenging in practice due to the complexities in humanoid perception and control, lingering physical gaps between humanoids and humans in morphologies and actuation, and lack of a data pipeline for humanoids to learn autonomous skills from egocentric vision. In this paper, we introduce a full-stack system for humanoids to learn motion and autonomous skills from human data. We first train a low-level policy in simulation via reinforcement learning using existing 40-hour human motion datasets. This policy transfers to the real world and allows humanoid robots to follow human body and hand motion in real time using only a RGB camera, i.e. shadowing. Through shadowing, human operators can teleoperate humanoids to collect whole-body data for learning different tasks in the real world. Using the data collected, we then perform supervised behavior cloning to train skill policies using egocentric vision, allowing humanoids to complete different tasks autonomously by imitating human skills. We demonstrate the system on our customized 33-DoF 180cm humanoid, autonomously completing tasks such as wearing a shoe to stand up and walk, unloading objects from warehouse racks, folding a sweatshirt, rearranging objects, typing, and greeting another robot with 60-100% success rates using up to 40 demonstrations. Project website: https://humanoid-ai.github.io/
Cited by
- From Generated Human Videos to Physically Plausible Robot Trajectories
- Learning Reusable Hybrid Motion Priors for Humanoid Locomotion from Motion Imitation
- EGM: Efficiently Learning General Motion Tracking Policy for High Dynamic Humanoid Whole-Body Control
- MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
- M4Human: A Large-Scale Multimodal mmWave Radar Benchmark for Human Mesh Reconstruction
- WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
- Toward Seamless Physical Human-Humanoid Interaction: Insights from Control, Intent, and Modeling with a Vision for What Comes Next
- Multi-Domain Motion Embedding: Expressive Real-Time Mimicry for Legged Robots
- Coordinated Humanoid Manipulation with Choice Policies
- HumanoidExo: Scalable Whole-Body Humanoid Manipulation via Wearable Exoskeleton
- Kinematics-Aware Multi-Policy Reinforcement Learning for Force-Capable Humanoid Loco-Manipulation
- SKEL-CF: Coarse-to-Fine Biomechanical Skeleton and Surface Mesh Recovery
- HAFO: A Force-Adaptive Control Framework for Humanoid Robots in Intense Interaction Environments
- SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control
- Structured Imitation Learning of Interactive Policies through Inverse Games
- Intuitive Programming, Adaptive Task Planning, and Dynamic Role Allocation in Human-Robot Collaboration
- Sim-to-Real Transfer in Deep Reinforcement Learning for Bipedal Locomotion
- GentleHumanoid: Learning Upper-body Compliance for Contact-rich Human and Object Interaction
- TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System
- Thor: Towards Human-Level Whole-Body Reactions for Intense Contact-Rich Environments
- PHUMA: Physically-Grounded Humanoid Locomotion Dataset
- Embracing Evolution: A Call for Body-Control Co-Design in Embodied Humanoid Robot
- One-shot Humanoid Whole-body Motion Learning
- HRT1: One-Shot Human-to-Robot Trajectory Transfer for Mobile Manipulation
- MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation
- SoftMimic: Learning Compliant Whole-body Control from Examples
- RAPID Hand Prototype: Design of an Affordable, Fully-Actuated Biomimetic Hand for Dexterous Teleoperation
- Towards Adaptable Humanoid Control via Adaptive Motion Tracking
- From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
- DemoHLM: From One Demonstration to Generalizable Humanoid Loco-Manipulation
- PhysHSI: Towards a Real-World Generalizable and Natural Humanoid-Scene Interaction System
- Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation
- Towards Proprioception-Aware Embodied Planning for Dual-Arm Humanoid Robots
- ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
- Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking
- ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning
- OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
- CoTaP: Compliant Task Pipeline and Reinforcement Learning of Its Controller with Compliance Modulation
- RobotDancing: Residual-Action Reinforcement Learning Enables Robust Long-Horizon Humanoid Motion Tracking
- HL-IK: A Lightweight Implementation of Human-Like Inverse Kinematics in Humanoid Arms
- EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data
- REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement Learning
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- KungfuBot2: Learning Versatile Motion Skills for Humanoid Whole-Body Control
- A Scalable Whole-body Motion Transfer via Implicit Kinodynamic Motion Retargeting
- DreamControl: Human-Inspired Whole-Body Humanoid Control for Scene Interaction via Guided Diffusion
- Track Any Motions under Any Disturbances
- StageACT: Stage-Conditioned Imitation for Robust Humanoid Door Opening
- Embracing Bulky Objects with Humanoid Robots: Whole-Body Manipulation with Reinforcement Learning
- EMMA: Scaling Mobile Manipulation via Egocentric Human Data
- HuBE: Cross-Embodiment Human-like Behavior Execution for Humanoid Robots
- Deep Sensorimotor Control by Imitating Predictive Models of Human Motion
- Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids
- Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse Reward
- MASH: Cooperative-Heterogeneous Multi-Agent Reinforcement Learning for Single Humanoid Robot Locomotion
- GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation
- Whole-Body Bilateral Teleoperation with Multi-Stage Object Parameter Estimation for Wheeled Humanoid Locomanipulation
- BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion
- Hand-Eye Autonomous Delivery: Learning Humanoid Navigation, Locomotion and Reaching
- TOP: Time Optimization Policy for Stable and Accurate Standing Manipulation with Humanoid Robots
- EMP: Executable Motion Prior for Humanoid Robot Standing Upper-body Motion Imitation
- Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots
- Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
- SkillBlender: Towards Versatile Humanoid Whole-Body Loco-Manipulation via Skill Blending
- CLONE: Closed-Loop Whole-Body Humanoid Teleoperation for Long-Horizon Tasks
- Robot Drummer: Learning Rhythmic Skills for Humanoid Drumming
- UniTracker: Learning Universal Whole-Body Motion Tracker for Humanoid Robots
- ULC: A Unified and Fine-Grained Controller for Humanoid Loco-Manipulation
- LOVON: Legged Open-Vocabulary Object Navigator
- Learning to Evaluate Autonomous Behaviour in Human-Robot Interaction
- Feature-Based vs. GAN-Based Learning from Demonstrations: When and Why
- Kinodynamic Motion Retargeting for Humanoid Locomotion via Multi-Contact Whole-Body Trajectory Optimization
- Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
- A Survey: Learning Embodied Intelligence from Physical Simulators and World Models
- ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation
- Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation
- Versatile Loco-Manipulation through Flexible Interlimb Coordination
- DualTHOR: A Dual-Arm Humanoid Simulation Platform for Contingency-Aware Planning
- Human2LocoMan: Learning Versatile Quadrupedal Manipulation with Human Pretraining
- TACT: Humanoid Whole-body Contact Manipulation through Deep Imitation Learning with Tactile Modality
- GMT: General Motion Tracking for Humanoid Whole-Body Control
- LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction
- KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills
- From Experts to a Generalist: Toward General Whole-Body Control for Humanoid Robots
- HoMeR: Learning In-the-Wild Mobile Manipulation via Hybrid Imitation and Whole-Body Control
- PyRoki: A Modular Toolkit for Robot Kinematic Optimization
- SignBot: Learning Human-to-Humanoid Sign Language Interaction
- Learning coordinated badminton skills for legged manipulators
- TWIST: Teleoperated Whole-Body Imitation System
- SMAP: Self-supervised Motion Adaptation for Physically Plausible Humanoid Whole-body Control
- DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion
- Teleopit: A Full-Embodiment Humanoid Teleoperation System
- Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
- Bootstrapping Imitation Learning for Long-horizon Manipulation via Hierarchical Data Collection Space
- Object-Focus Actor for Data-efficient Robot Generalization Dexterous Manipulation
- Dribble Master: Learning Agile Humanoid Dribbling through Legged Locomotion
- X2C: A Dataset Featuring Nuanced Facial Expressions for Realistic Humanoid Imitation
- Unleashing Humanoid Reaching Potential via Real-world-Ready Skill Space
- HuB: Learning Extreme Humanoid Balance
- JAEGER: Dual-Level Humanoid Whole-Body Controller
- FALCON: Learning Force-Adaptive Humanoid Loco-Manipulation
- WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation
- A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
- Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching
- LangWBC: Language-directed Humanoid Whole-Body Control via End-to-end Learning
- Learning Action Priors for Cross-embodiment Robot Manipulation
- OmniXtreme: Breaking the Generality Barrier in High-Dynamic Humanoid Control
- HumanX: Toward Agile and Generalizable Humanoid Interaction Skills from Human Videos
- STDArm: Transferring Visuomotor Policies From Static Data Training to Dynamic Robot Manipulation
- MotionPRO: Exploring the Role of Pressure in Human MoCap and Beyond
- Physically Consistent Humanoid Loco-Manipulation using Latent Diffusion Models
- Adversarial Locomotion and Motion Imitation for Humanoid Policy Learning
- Spectral Normalization for Lipschitz-Constrained Policies on Learning Humanoid Locomotion
Related