2019/11/17 by Abhinav Jain, Jain, Abhinav, Frank Dellaert +1
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Robot Manipulation and Learning #Robotics and Sensor-Based Localization
paper · pdf · doi:10.48550/arxiv.1911.07347
openalex publication_date 2019/11/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Pose estimation is a vital step in many robotics and perception tasks such as robotic manipulation, autonomous vehicle navigation, etc. Current state-of-the-art pose estimation methods rely on deep neural networks with complicated structures and long inference times. While highly robust, they require computing power often unavailable on mobile robots. We propose a CNN-based pose refinement system which takes a coarsely estimated 3D pose from a computationally cheaper algorithm along with a bounding box image of the object, and returns a highly refined pose. Our experiments on the YCB-Video dataset show that our system can refine 3D poses to an extremely high precision with minimal training data.