CLIPort: What and Where Pathways for Robotic Manipulation
2021/09/24 by Shridhar, Mohit, Manuelli, Lucas, Fox, Dieter · 83 citations
#Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2109.12098
Abstract
How can we imbue robots with the ability to manipulate objects precisely but also to reason about them in terms of abstract concepts? Recent works in manipulation have shown that end-to-end networks can learn dexterous skills that require precise spatial reasoning, but these methods often fail to generalize to new goals or quickly learn transferable concepts across tasks. In parallel, there has been great progress in learning generalizable semantic representations for vision and language by training on large-scale internet data, however these representations lack the spatial understanding necessary for fine-grained manipulation. To this end, we propose a framework that combines the best of both worlds: a two-stream architecture with semantic and spatial pathways for vision-based manipulation. Specifically, we present CLIPort, a language-conditioned imitation-learning agent that combines the broad semantic understanding (what) of CLIP [1] with the spatial precision (where) of Transporter [2]. Our end-to-end framework is capable of solving a variety of language-specified tabletop tasks from packing unseen objects to folding cloths, all without any explicit representations of object poses, instance segmentations, memory, symbolic states, or syntactic structures. Experiments in simulated and real-world settings show that our approach is data efficient in few-shot settings and generalizes effectively to seen and unseen semantic concepts. We even learn one multi-task policy for 10 simulated and 9 real-world tasks that is better or comparable to single-task policies.
Cited by
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation
- Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
- AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning
- VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation
- TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models
- Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
- OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning
- Multi-Robot Motion Planning from Vision and Language using Heat-Inspired Diffusion
- SAGA: Open-World Mobile Manipulation via Structured Affordance Grounding
- Architecting Large Action Models for Human-in-the-Loop Intelligent Robots
- Towards Logic-Aware Manipulation: A Knowledge Primitive for VLM-Based Assistants in Smart Manufacturing
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination
- VAT: Vision Action Transformer by Unlocking Full Representation of ViT
- NOIR 2.0: Neural Signal Operated Intelligent Robots for Everyday Activities
- ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation
- Arcadia: Toward a Full-Lifecycle Framework for Embodied Lifelong Learning
- SFHand: A Streaming Framework for Language-guided 3D Hand Forecasting and Embodied Manipulation
- EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
- Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
- C2F-Space: Coarse-to-Fine Space Grounding for Spatial Instructions using Vision-Language Models
- Eq.Bot: Enhance Robotic Manipulation Learning via Group Equivariant Canonicalization
- Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception
- EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
- Simulating the Visual World with Artificial Intelligence: A Roadmap
- VLAD-Grasp: Zero-shot Grasp Detection via Vision-Language Models
- Lite VLA: Efficient Vision-Language-Action Control on CPU-Bound Edge Robots
- LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation
- Text to Robotic Assembly of Multi Component Objects using 3D Generative AI and Vision Language Models
- To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning
- PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI
- Generalizable Hierarchical Skill Learning via Object-Centric Representation
- Exploring Conditions for Diffusion models in Robotic Control
- A Compositional Paradigm for Foundation Models: Towards Smarter Robotic Agents
- Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Avi: Action from Volumetric Inference
- ARRC: Advanced Reasoning Robot Control - Knowledge-Driven Autonomous Manipulation Using Retrieval-Augmented Generation
- LangGrasp: Leveraging Fine-Tuned LLMs for Language Interactive Robot Grasping with Ambiguous Instructions
- Mash, Spread, Slice! Learning to Manipulate Object States via Visual Spatial Progress
- LAGEA: Language Guided Embodied Agents for Robotic Manipulation
- IROSA: Interactive Robot Skill Adaptation Using Natural Language
- Hybrid Diffusion for Simultaneous Symbolic and Continuous Planning
- EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transfer
- Look as You Leap: Planning Simultaneous Motion and Perception for High-DOF Robots
- Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation
- SINGER: An Onboard Generalist Vision-Language Navigation Policy for Drones
- FUNCanon: Learning Pose-Aware Action Primitives via Functional Object Canonicalization for Generalizable Robotic Manipulation
- 3D Flow Diffusion Policy: Visuomotor Policy Learning via Generating Flow in 3D Space
- No Need for Real 3D: Fusing 2D Vision with Pseudo 3D Representations for Robotic Manipulation Learning
- GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation
- Force-Modulated Visual Policy for Robot-Assisted Dressing with Arm Motions
- Pre-trained Visual Representations Generalize Where it Matters in Model-Based Reinforcement Learning
- Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
- TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making
- Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning
- Graph-Fused Vision-Language-Action for Policy Reasoning in Multi-Arm Robotic Manipulation
- Imitation Learning Based on Disentangled Representation Learning of Behavioral Characteristics
- Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception
- ConceptBot: Enhancing Robot's Autonomy through Task Decomposition with Large Language Models and Knowledge Graph
- CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
- HyperTASR: Hypernetwork-Driven Task-Aware Scene Representations for Robust Manipulation
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Human Centric General Physical Intelligence for Agile Manufacturing Automation
- Reinforcing Video Reasoning Segmentation to Think Before It Segments
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT
- Rational Inverse Reasoning: Few-Shot Imitation by Inferring Intent through Planning
- AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Mixed-Initiative Dialog for Human-Robot Collaborative Manipulation
- Information-Theoretic Graph Fusion with Vision-Language-Action Model for Policy Reasoning and Dual Robotic Control
- ASkDAgger: Active Skill-level Data Aggregation for Interactive Imitation Learning
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
- Analyzing the Impact of Multimodal Perception on Sample Complexity and Optimization Landscapes in Imitation Learning
- Improving Generalization of Language-Conditioned Robot Manipulation
- RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
Related