2020/07/21 by He Chen, Pengfei Guo, Chen, He +8 · 3 citations
Computer Science · Mathematics · #3D pose estimation #Advanced Vision and Imaging #Ambiguity #Artificial intelligence #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Computer vision #Epipolar geometry #Estimator #FOS: Computer and information sciences #Feature (linguistics) #Human Pose and Action Recognition #Image (mathematics) #Matching (statistics) #Mathematics #Pattern recognition (psychology) #Pose #Robustness (evolution) #Segmentation #Video Surveillance and Tracking Methods #cs.CV
paper · pdf · doi:10.48550/arxiv.2007.10986
published in arXiv (Cornell University) (Cornell University)
arxiv created 2020/07/21 · openalex publication_date 2020/07/21 · arxiv updated 2020/07/22 · openalex created_date 2020/07/29 · openalex updated_date 2026/08/06
Epipolar constraints are at the core of feature matching and depth estimation in current multi-person multi-camera 3D human pose estimation methods. Despite the satisfactory performance of this formulation in sparser crowd scenes, its effectiveness is frequently challenged under denser crowd circumstances mainly due to two sources of ambiguity. The first is the mismatch of human joints resulting from the simple cues provided by the Euclidean distances between joints and epipolar lines. The second is the lack of robustness from the naive formulation of the problem as a least squares minimization. In this paper, we depart from the multi-person 3D pose estimation formulation, and instead reformulate it as crowd pose estimation. Our method consists of two key components: a graph model for fast cross-view matching, and a maximum a posteriori (MAP) estimator for the reconstruction of the 3D human poses. We demonstrate the effectiveness and superiority of our proposed method on four benchmark datasets.