HumanPlus: Humanoid Shadowing and Imitation from Humans
2024/06/15 by Zipeng Fu, Fu, Zipeng, Qingqing Zhao +7 · 68 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Human Pose and Action Recognition #Machine Learning (cs.LG) #Robotics (cs.RO) #Systems and Control (eess.SY) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2406.10454
openalex publication_date 2024/06/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
One of the key arguments for building robots that have similar form factors to human beings is that we can leverage the massive human data for training. Yet, doing so has remained challenging in practice due to the complexities in humanoid perception and control, lingering physical gaps between humanoids and humans in morphologies and actuation, and lack of a data pipeline for humanoids to learn autonomous skills from egocentric vision. In this paper, we introduce a full-stack system for humanoids to learn motion and autonomous skills from human data. We first train a low-level policy in simulation via reinforcement learning using existing 40-hour human motion datasets. This policy transfers to the real world and allows humanoid robots to follow human body and hand motion in real time using only a RGB camera, i.e. shadowing. Through shadowing, human operators can teleoperate humanoids to collect whole-body data for learning different tasks in the real world. Using the data collected, we then perform supervised behavior cloning to train skill policies using egocentric vision, allowing humanoids to complete different tasks autonomously by imitating human skills. We demonstrate the system on our customized 33-DoF 180cm humanoid, autonomously completing tasks such as wearing a shoe to stand up and walk, unloading objects from warehouse racks, folding a sweatshirt, rearranging objects, typing, and greeting another robot with 60-100% success rates using up to 40 demonstrations. Project website: https://humanoid-ai.github.io/
Cited by
- From Generated Human Videos to Physically Plausible Robot Trajectories
- Learning Reusable Hybrid Motion Priors for Humanoid Locomotion from Motion Imitation
- EGM: Efficiently Learning General Motion Tracking Policy for High Dynamic Humanoid Whole-Body Control
- MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
- M4Human: A Large-Scale Multimodal mmWave Radar Benchmark for Human Mesh Reconstruction
- WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
- Toward Seamless Physical Human-Humanoid Interaction: Insights from Control, Intent, and Modeling with a Vision for What Comes Next
- Multi-Domain Motion Embedding: Expressive Real-Time Mimicry for Legged Robots
- Coordinated Humanoid Manipulation with Choice Policies
- HumanoidExo: Scalable Whole-Body Humanoid Manipulation via Wearable Exoskeleton
- Kinematics-Aware Multi-Policy Reinforcement Learning for Force-Capable Humanoid Loco-Manipulation
- SKEL-CF: Coarse-to-Fine Biomechanical Skeleton and Surface Mesh Recovery
- HAFO: A Force-Adaptive Control Framework for Humanoid Robots in Intense Interaction Environments
- SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control
- Structured Imitation Learning of Interactive Policies through Inverse Games
- Intuitive Programming, Adaptive Task Planning, and Dynamic Role Allocation in Human-Robot Collaboration
- Sim-to-Real Transfer in Deep Reinforcement Learning for Bipedal Locomotion
- GentleHumanoid: Learning Upper-body Compliance for Contact-rich Human and Object Interaction
- TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System
- Thor: Towards Human-Level Whole-Body Reactions for Intense Contact-Rich Environments
- PHUMA: Physically-Grounded Humanoid Locomotion Dataset
- Embracing Evolution: A Call for Body-Control Co-Design in Embodied Humanoid Robot
- One-shot Humanoid Whole-body Motion Learning
- HRT1: One-Shot Human-to-Robot Trajectory Transfer for Mobile Manipulation
- MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation
- SoftMimic: Learning Compliant Whole-body Control from Examples
- RAPID Hand Prototype: Design of an Affordable, Fully-Actuated Biomimetic Hand for Dexterous Teleoperation
- Towards Adaptable Humanoid Control via Adaptive Motion Tracking
- From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
- DemoHLM: From One Demonstration to Generalizable Humanoid Loco-Manipulation
- PhysHSI: Towards a Real-World Generalizable and Natural Humanoid-Scene Interaction System
- Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation
- Towards Proprioception-Aware Embodied Planning for Dual-Arm Humanoid Robots
- ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
- Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking
- ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning
- OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
- CoTaP: Compliant Task Pipeline and Reinforcement Learning of Its Controller with Compliance Modulation
- RobotDancing: Residual-Action Reinforcement Learning Enables Robust Long-Horizon Humanoid Motion Tracking
- HL-IK: A Lightweight Implementation of Human-Like Inverse Kinematics in Humanoid Arms
- EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data
- REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement Learning
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- KungfuBot2: Learning Versatile Motion Skills for Humanoid Whole-Body Control
- A Scalable Whole-body Motion Transfer via Implicit Kinodynamic Motion Retargeting
- DreamControl: Human-Inspired Whole-Body Humanoid Control for Scene Interaction via Guided Diffusion
- Track Any Motions under Any Disturbances
- StageACT: Stage-Conditioned Imitation for Robust Humanoid Door Opening
- Embracing Bulky Objects with Humanoid Robots: Whole-Body Manipulation with Reinforcement Learning
- EMMA: Scaling Mobile Manipulation via Egocentric Human Data
- HuBE: Cross-Embodiment Human-like Behavior Execution for Humanoid Robots
- Deep Sensorimotor Control by Imitating Predictive Models of Human Motion
- Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids
- Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse Reward
- MASH: Cooperative-Heterogeneous Multi-Agent Reinforcement Learning for Single Humanoid Robot Locomotion
- GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation
- Whole-Body Bilateral Teleoperation with Multi-Stage Object Parameter Estimation for Wheeled Humanoid Locomanipulation
- BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion
- Hand-Eye Autonomous Delivery: Learning Humanoid Navigation, Locomotion and Reaching
- TOP: Time Optimization Policy for Stable and Accurate Standing Manipulation with Humanoid Robots
- EMP: Executable Motion Prior for Humanoid Robot Standing Upper-body Motion Imitation
- Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots
- Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
- Robot Drummer: Learning Rhythmic Skills for Humanoid Drumming
- UniTracker: Learning Universal Whole-Body Motion Tracker for Humanoid Robots
- ULC: A Unified and Fine-Grained Controller for Humanoid Loco-Manipulation
- LOVON: Legged Open-Vocabulary Object Navigator
- Learning to Evaluate Autonomous Behaviour in Human-Robot Interaction
Related