2016/05/10 by Abhilash Srikantha, Srikantha, Abhilash, Jüergen Gall +1
Computer Science · Engineering · #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Reinforcement Learning in Robotics #Robot Manipulation and Learning
paper · pdf · doi:10.48550/arxiv.1605.02964
openalex publication_date 2016/05/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Localizing functional regions of objects or affordances is an important aspect of scene understanding. In this work, we cast the problem of affordance segmentation as that of semantic image segmentation. In order to explore various levels of supervision, we introduce a pixel-annotated affordance dataset of 3090 images containing 9916 object instances with rich contextual information in terms of human-object interactions. We use a deep convolutional neural network within an expectation maximization framework to take advantage of weakly labeled data like image level annotations or keypoint annotations. We show that a further reduction in supervision is possible with a minimal loss in performance when human pose is used as context.