vix.ing · top · new · best · stats · spec

4D Human Body Capture from Egocentric Video via 3D Scene Grounding

2020/11/26 by Miao Liu, Dexin Yang, Liu, Miao +9
Computer Science · Engineering · #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Motion and Animation #Human Pose and Action Recognition #cs.CV

paper · pdf · doi:10.48550/arxiv.2011.13341

openalex publication_date 2020/11/26 · arxiv created 2021/10/15 · arxiv updated 2021/10/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We introduce a novel task of reconstructing a time series of second-person 3D human body meshes from monocular egocentric videos. The unique viewpoint and rapid embodied camera motion of egocentric videos raise additional technical barriers for human body capture. To address those challenges, we propose a simple yet effective optimization-based approach that leverages 2D observations of the entire video sequence and human-scene interaction constraint to estimate second-person human poses, shapes, and global motion that are grounded on the 3D environment captured from the egocentric view. We conduct detailed ablation studies to validate our design choice. Moreover, we compare our method with the previous state-of-the-art method on human motion capture from monocular video, and show that our method estimates more accurate human-body poses and shapes under the challenging egocentric setting. In addition, we demonstrate that our approach produces more realistic human-scene interaction.

Citations

Related