vix.ing · top · new · best · stats · spec

Compose by Focus: Scene Graph-based Atomic Skills

2025/09/19 by Qi Han, Qi, Han, Changhe Chen +3
Decision Sciences · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Robotics (cs.RO) #Scientific Computing and Data Management

paper · pdf · doi:10.48550/arxiv.2509.16053

openalex publication_date 2025/09/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

A key requirement for generalist robots is compositional generalization - the ability to combine atomic skills to solve complex, long-horizon tasks. While prior work has primarily focused on synthesizing a planner that sequences pre-learned skills, robust execution of the individual skills themselves remains challenging, as visuomotor policies often fail under distribution shifts induced by scene composition. To address this, we introduce a scene graph-based representation that focuses on task-relevant objects and relations, thereby mitigating sensitivity to irrelevant variation. Building on this idea, we develop a scene-graph skill learning framework that integrates graph neural networks with diffusion-based imitation learning, and further combine "focused" scene-graph skills with a vision-language model (VLM) based task planner. Experiments in both simulation and real-world manipulation tasks demonstrate substantially higher success rates than state-of-the-art baselines, highlighting improved robustness and compositional generalization in long-horizon tasks.

Citations

Related