vix.ing · top · new · best · stats · spec

PersonLab: Person Pose Estimation and Instance Segmentation with a\n Bottom-Up, Part-Based, Geometric Embedding Model

2018/03/22 by George Papandreou, Papandreou, George, Tyler Zhu +9 · 2 citations
Computer Science · #Human Pose and Action Recognition #Video Surveillance and Tracking Methods #Anomaly Detection Techniques and Applications

paper · pdf · doi:10.48550/arxiv.1803.08225

Abstract

We present a box-free bottom-up approach for the tasks of pose estimation and\ninstance segmentation of people in multi-person images using an efficient\nsingle-shot model. The proposed PersonLab model tackles both semantic-level\nreasoning and object-part associations using part-based modeling. Our model\nemploys a convolutional network which learns to detect individual keypoints and\npredict their relative displacements, allowing us to group keypoints into\nperson pose instances. Further, we propose a part-induced geometric embedding\ndescriptor which allows us to associate semantic person pixels with their\ncorresponding person instance, delivering instance-level person segmentations.\nOur system is based on a fully-convolutional architecture and allows for\nefficient inference, with runtime essentially independent of the number of\npeople present in the scene. Trained on COCO data alone, our system achieves\nCOCO test-dev keypoint average precision of 0.665 using single-scale inference\nand 0.687 using multi-scale inference, significantly outperforming all previous\nbottom-up pose estimation systems. We are also the first bottom-up method to\nreport competitive results for the person class in the COCO instance\nsegmentation task, achieving a person category average precision of 0.417.\n

Cited by

Related