RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
2025/02/28 by Ji, Yuheng, Tan, Huajie, Shi, Jiayu +14 · 66 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2502.21257
Abstract
Recent advancements in Multimodal Large Language Models (MLLMs) have shown remarkable capabilities across various multimodal contexts. However, their application in robotic scenarios, particularly for long-horizon manipulation tasks, reveals significant limitations. These limitations arise from the current MLLMs lacking three essential robotic brain capabilities: Planning Capability, which involves decomposing complex manipulation instructions into manageable sub-tasks; Affordance Perception, the ability to recognize and interpret the affordances of interactive objects; and Trajectory Prediction, the foresight to anticipate the complete manipulation trajectory necessary for successful execution. To enhance the robotic brain's core capabilities from abstract to concrete, we introduce ShareRobot, a high-quality heterogeneous dataset that labels multi-dimensional information such as task planning, object affordance, and end-effector trajectory. ShareRobot's diversity and accuracy have been meticulously refined by three human annotators. Building on this dataset, we developed RoboBrain, an MLLM-based model that combines robotic and general multi-modal data, utilizes a multi-stage training strategy, and incorporates long videos and high-resolution images to improve its robotic manipulation capabilities. Extensive experiments demonstrate that RoboBrain achieves state-of-the-art performance across various robotic tasks, highlighting its potential to advance robotic brain capabilities.
Cited by
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation
- Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
- Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
- SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment
- Intra-Class Probabilistic Embeddings for Uncertainty Estimation in Vision-Language Models
- Towards Cross-View Point Correspondence in Vision-Language Models
- SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL
- OmniDexVLG: Learning Dexterous Grasp Generation from Vision Language Model-Guided Grasp Semantics, Taxonomy and Functional Affordance
- Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
- G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
- RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation
- BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
- FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
- RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
- iFlyBot-VLM Technical Report
- GraspView: Active Perception Scoring and Best-View Optimization for Robotic Grasping in Cluttered Environments
- RoboOS-NeXT: A Unified Memory-based Framework for Lifelong, Scalable, and Robust Multi-Robot Collaboration
- Pelican-VL 1.0: A Foundation Brain Model for Embodied Intelligence
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning
- Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
- RoboGPT-R1: Enhancing Robot Planning with Reinforcement Learning
- From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
- EmboMatrix: A Scalable Training-Ground for Embodied Decision-Making
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
- RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action Model
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
- TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
- Contrastive Representation Regularization for Vision-Language-Action Models
- MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
- Reinforced Embodied Planning with Verifiable Reward for Real-World Robotic Manipulation
- Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
- Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
- Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- RoboSeek: You Need to Interact with Your Objects
- Video-to-BT: Generating Reactive Behavior Trees from Human Demonstration Videos for Robotic Assembly
- Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- Planning with Reasoning using Vision Language World Model
- RynnEC: Bringing MLLMs into Embodied World
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- NavA3: Understanding Any Instruction, Navigating Anywhere, Finding Anything
- Diffusion Models are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning
- Vision Language Action Models in Robotic Manipulation: A Systematic Review
- VIKI-R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning
- Synergistic Prompting for Robust Visual Recognition with Missing Modalities
- Training-free Generation of Temporally Consistent Rewards from VLMs
- HumanoidGen: Data Generation for Bimanual Dexterous Manipulation via LLM Reasoning
- Parallels Between VLA Model Post-Training and Human Motor Learning: Progress, Challenges, and Trends
- FrankenBot: Brain-Morphic Modular Orchestration for Robotic Manipulation with Vision-Language Models
- VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
- Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
Related