Human3R: Everyone Everywhere All at Once
2025/10/07 by Yue Chen, Chen, Yue, Xingyu Chen +9 · 4 citations
Computer Science · Medicine · #Context-Aware Activity Recognition Systems #COVID-19 diagnosis using AI
paper · pdf · doi:10.48550/arxiv.2510.06219
Abstract
We present Human3R, a unified, feed-forward framework for online 4D human-scene reconstruction, in the world frame, from casually captured monocular videos. Unlike previous approaches that rely on multi-stage pipelines, iterative contact-aware refinement between humans and scenes, and heavy dependencies, e.g., human detection, depth estimation, and SLAM pre-processing, Human3R jointly recovers global multi-person SMPL-X bodies ("everyone"), dense 3D scene ("everywhere"), and camera trajectories in a single forward pass ("all-at-once"). Our method builds upon the 4D online reconstruction model CUT3R, and uses parameter-efficient visual prompt tuning, to strive to preserve CUT3R's rich spatiotemporal priors, while enabling direct readout of multiple SMPL-X bodies. Human3R is a unified model that eliminates heavy dependencies and iterative refinement. After being trained on the relatively small-scale synthetic dataset BEDLAM for just one day on one GPU, it achieves superior performance with remarkable efficiency: it reconstructs multiple humans in a one-shot manner, along with 3D scenes, in one stage, in real-time (15 FPS) with a low memory footprint (8 GB). Extensive experiments demonstrate that Human3R delivers state-of-the-art or competitive performance across tasks, including global human motion estimation, local human mesh recovery, video depth estimation, and camera pose estimation, with a single unified model. We hope that Human3R will serve as a simple yet strong baseline, which can be easily adapted for downstream applications. Code, models and 4D interactive demos are available at https://fanegg.github.io/Human3R/.
Citations
- TTT3R: 3D Reconstruction as Test-Time Training
- HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction
- LONG3R: Long Sequence Streaming 3D Reconstruction
- Streaming 4D Visual Geometry Transformer
- Understanding and Improving Length Generalization in Recurrent Models
- Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
- Visual Imitation Enables Contextual Humanoid Control
- PromptHMR: Promptable Human Mesh Recovery
- Easi3R: Estimating Disentangled Motion from DUSt3R Without Training
- VGGT: Visual Geometry Grounded Transformer
- MUSt3R: Multi-view Network for Stereo 3D Reconstruction
- Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
- Continuous 3D Perception Model with Persistent State
- Joint Optimization for 4D Human-Scene Reconstruction in the Wild
- Humans as a Calibration Pattern: Dynamic 3D Scene Reconstruction from Unsynchronized and Uncalibrated Videos
- Reconstructing People, Places, and Cameras
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- Feat2GS: Probing Visual Foundation Models with Gaussian Splatting
- MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
- DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruction
- CameraHMR: Aligning People with Perspective
- MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
- MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion
- COIN: Control-Inpainting Diffusion Prior for Human and Camera Motion Estimation
- 3D Reconstruction with Spatial Memory
- SAM 2: Segment Anything in Images and Videos
- Neural Localizer Fields for Continuous 3D Human Pose and Shape Estimation
- Learning to (Learn at Test Time): RNNs with Expressive Hidden States
- Grounding Image Matching in 3D with MASt3R
- Synergistic Global-space Camera and Human Reconstruction from Videos
- TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation
- Scene Coordinate Reconstruction: Posing of Image Collections via Incremental Learning of a Relocalizer
- TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos
- Multi-HMR: Multi-Person Whole-Body Human Mesh Recovery in a Single Shot
- Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
- DUSt3R: Geometric 3D Vision Made Easy
- WHAM: Reconstructing World-grounded Humans with Accurate 3D Motion
- ChatPose: Chatting about 3D Human Pose
- PACE: Human and Camera Motion Estimation from in-the-wild Videos
- Tracking Anything with Decoupled Video Segmentation
- EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild
- TRACE: 5D Temporal Regression of Avatars with Dynamic Cameras in 3D Environments
- Humans in 4D: Reconstructing and Tracking Humans with Transformers
- DINOv2: Learning Robust Visual Features without Supervision
- HybrIK-X: Hybrid Analytical-Neural Inverse Kinematics for Whole-body Mesh Recovery
- Micrograph segmentations for DDEVD
- Segment Anything
- SLOPER4D: A Scene-Aware Dataset for Global 4D Human Pose Estimation in Urban Environments
- Decoupling Human and Camera Motion from Videos in the Wild
- ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth
- CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical Flow
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion
- The One Where They Reconstructed 3D Humans and Environments in TV Shows
- ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation
- Exploring Plain Vision Transformer Backbones for Object Detection
- Visual Prompt Tuning
- Putting People in their Place: Monocular Regression of 3D People in Depth
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic Cameras
- SPEC: Seeing People in the Wild with an Estimated Camera
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras
- Collaborative Regression of Expressive Bodies using Moderation
- HuMoR: 3D Human Motion Model for Robust Pose Estimation
- Emerging Properties in Self-Supervised Vision Transformers
- Emerging Properties in Self-Supervised Vision Transformers
- PyMAF: 3D Human Pose and Shape Regression with Pyramidal Mesh Alignment Feedback Loop
- Vision Transformers for Dense Prediction
- HybrIK: A Hybrid Analytical-Neural Inverse Kinematics Solution for 3D Human Pose and Shape Estimation
- Beyond Static Features for Temporally Consistent 3D Human Pose and Shape from a Video
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Monocular, One-stage, Regression of Multiple 3D People
- Monocular Expressive Body Regression through Body-Driven Attention
- Exemplar Fine-Tuning for 3D Human Model Fitting Towards In-the-Wild 3D Human Pose Estimation
- VIBE: Video Inference for Human Body Pose and Shape Estimation
- Learning to Reconstruct 3D Human Pose and Shape via Model-fitting in the\n Loop
- Resolving 3D Human Pose Ambiguities with 3D Scene Constraints
- ReFusion: 3D Reconstruction in Dynamic Environments for RGB-D Cameras\n Exploiting Residuals
- Objects as Points
- Expressive Body Capture: 3D Hands, Face, and Body from a Single Image
- Learning 3D Human Dynamics from Video
- Neural Body Fitting: Unifying Deep Learning and Model-Based Human Pose and Shape Estimation
- DensePose: Dense Human Pose Estimation In The Wild
- End-to-end Recovery of Human Shape and Pose
- Attention Is All You Need
- DSAC - Differentiable RANSAC for Camera Localization
- Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image
- Gaussian Error Linear Units (GELUs)
- Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network
Cited by
Related