vix.ing · top · new · best · stats

Affordance-Guided Reinforcement Learning via Visual Prompting

2024/07/14 by Olivia Y. Lee, Annie Xie, Lee, Olivia Y. +7 · 14 citations
Computer Science · Psychology · #Affordance #Artificial Intelligence (cs.AI) #Artificial intelligence #Cognitive psychology #Cognitive science #Computer science #FOS: Computer and information sciences #Human–computer interaction #Machine Learning (cs.LG) #Psychology #Reinforcement #Reinforcement Learning in Robotics #Reinforcement learning #Robotics (cs.RO) #Social psychology

paper · pdf · doi:10.48550/arxiv.2407.10341

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2024/07/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Robots equipped with reinforcement learning (RL) have the potential to learn a wide range of skills solely from a reward signal. However, obtaining a robust and dense reward signal for general manipulation tasks remains a challenge. Existing learning-based approaches require significant data, such as human demonstrations of success and failure, to learn task-specific reward functions. Recently, there is also a growing adoption of large multi-modal foundation models for robotics that can perform visual reasoning in physical contexts and generate coarse robot motions for manipulation tasks. Motivated by this range of capability, in this work, we present Keypoint-based Affordance Guidance for Improvements (KAGI), a method leveraging rewards shaped by vision-language models (VLMs) for autonomous RL. State-of-the-art VLMs have demonstrated impressive zero-shot reasoning about affordances through keypoints, and we use these to define dense rewards that guide autonomous robotic learning. On diverse real-world manipulation tasks specified by natural language descriptions, KAGI improves the sample efficiency of autonomous RL and enables successful task completion in 30K online fine-tuning steps. Additionally, we demonstrate the robustness of KAGI to reductions in the number of in-domain demonstrations used for pre-training, reaching similar performance in 45K online fine-tuning steps. Project website: https://sites.google.com/view/affordance-guided-rl

Cited by

Related