vix.ing · top · new · best · stats · spec

Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models

2023/12/07 by Ivan Kapelyukh, Yifei Ren, Kapelyukh, Ivan +5 · 3 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Robotics (cs.RO)

paper · pdf · doi:10.48550/arxiv.2312.04533

openalex publication_date 2023/12/07 · openalex created_date 2023/12/09 · openalex updated_date 2026/07/28

Abstract

We introduce Dream2Real, a robotics framework which integrates vision-language models (VLMs) trained on 2D data into a 3D object rearrangement pipeline. This is achieved by the robot autonomously constructing a 3D representation of the scene, where objects can be rearranged virtually and an image of the resulting arrangement rendered. These renders are evaluated by a VLM, so that the arrangement which best satisfies the user instruction is selected and recreated in the real world with pick-and-place. This enables language-conditioned rearrangement to be performed zero-shot, without needing to collect a training dataset of example arrangements. Results on a series of real-world tasks show that this framework is robust to distractors, controllable by language, capable of understanding complex multi-object relations, and readily applicable to both tabletop and 6-DoF rearrangement tasks.

Cited by

Related