vix.ing · top · new · best · stats · spec

Leveraging Photometric Consistency over Time for Sparsely Supervised\n Hand-Object Reconstruction

2020/04/28 by Yana Hasson, Hasson, Yana, Bugra Tekin +9 · 7 citations
Computer Science · #Human Pose and Action Recognition #Video Surveillance and Tracking Methods #Advanced Vision and Imaging

paper · pdf · doi:10.48550/arxiv.2004.13449

Abstract

Modeling hand-object manipulations is essential for understanding how humans\ninteract with their environment. While of practical importance, estimating the\npose of hands and objects during interactions is challenging due to the large\nmutual occlusions that occur during manipulation. Recent efforts have been\ndirected towards fully-supervised methods that require large amounts of labeled\ntraining samples. Collecting 3D ground-truth data for hand-object interactions,\nhowever, is costly, tedious, and error-prone. To overcome this challenge we\npresent a method to leverage photometric consistency across time when\nannotations are only available for a sparse subset of frames in a video. Our\nmodel is trained end-to-end on color images to jointly reconstruct hands and\nobjects in 3D by inferring their poses. Given our estimated reconstructions, we\ndifferentiably render the optical flow between pairs of adjacent images and use\nit within the network to warp one frame to another. We then apply a\nself-supervised photometric loss that relies on the visual consistency between\nnearby images. We achieve state-of-the-art results on 3D hand-object\nreconstruction benchmarks and demonstrate that our approach allows us to\nimprove the pose estimation accuracy by leveraging information from neighboring\nframes in low-data regimes.\n

Cited by

Related