RT-1: Robotics Transformer for Real-World Control at Scale
2022/12/13 by Anthony Brohan, Brohan, Anthony, Noah Brown +99 · 373 citations
Computer Science · #Advanced Neural Network Applications #Multimodal Machine Learning Applications #Machine Learning and Data Classification
paper · pdf · doi:10.48550/arxiv.2212.06817
Abstract
By transferring knowledge from large, diverse, task-agnostic datasets, modern machine learning models can solve specific downstream tasks either zero-shot or with small task-specific datasets to a high level of performance. While this capability has been demonstrated in other fields such as computer vision, natural language processing or speech recognition, it remains to be shown in robotics, where the generalization capabilities of the models are particularly critical due to the difficulty of collecting real-world robotic data. We argue that one of the keys to the success of such general robotic models lies with open-ended task-agnostic training, combined with high-capacity architectures that can absorb all of the diverse, robotic data. In this paper, we present a model class, dubbed Robotics Transformer, that exhibits promising scalable model properties. We verify our conclusions in a study of different model classes and their ability to generalize as a function of the data size, model size, and data diversity based on a large-scale data collection on real robots performing real-world tasks. The project's website and videos can be found at robotics-transformer1.github.io
Cited by
- Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
- Embodied Learning of Reward for Musculoskeletal Control with Vision Language Models
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- Autoregressive Flow Matching for Motion Prediction
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
- Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- N0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
- N0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
- WCM: World-Cognition Model for Generalizable Human-Robot Interaction
- DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
- The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation
- Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
- CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model
- Physical AI Governance: From Theory to Practice Across Life Cycle
- RoboCade: Gamifying Robot Data Collection
- LoLA: Long Horizon Latent Action Learning for General Robot Manipulation
- Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation
- REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation
- MaP-AVR: A Meta-Action Planner for Agents Leveraging Vision Language Models and Retrieval-Augmented Generation
- Point What You Mean: Visually Grounded Instruction Policy
- Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
- Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
- PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence
- VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation
- Large Video Planner Enables Generalizable Robot Control
- VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
- EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning
- Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
- SAGA: Open-World Mobile Manipulation via Structured Affordance Grounding
- Towards Logic-Aware Manipulation: A Knowledge Primitive for VLM-Based Assistants in Smart Manufacturing
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
- RoboNeuron: A Middle-Layer Infrastructure for Agent-Driven Orchestration in Embodied AI
- Openpi Comet: Competition Solution For 2025 BEHAVIOR Challenge
- Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning
- GLaD: Geometric Latent Distillation for Vision-Language-Action Models
- H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
- Masked Generative Policy for Robotic Control
- Astra: General Interactive World Model with Autoregressive Denoising
- Bridging Scale Discrepancies in Robotic Control via Language-Based Action Representations
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- Delay-Aware Diffusion Policy: Bridging the Observation-Execution Gap in Dynamic Tasks
- ESPADA: Execution Speedup via Semantics Aware Demonstration Data Downsampling for Imitation Learning
- See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
- Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge
- A Novel Multimodal RUL Framework for Remaining Useful Life Estimation with Layer-wise Explanations
- Stitch and Tell: A Structured Multimodal Data Augmentation Method for Spatial Understanding
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- Invariance Co-training for Robot Visual Generalization
- Correspondence-Oriented Imitation Learning: Flexible Visuomotor Control with 3D Conditioning
- HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
- STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models
- Object Reconstruction under Occlusion with Generative Priors and Contact-induced Constraints
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination
- Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
- PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
- VAT: Vision Action Transformer by Unlocking Full Representation of ViT
- What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
- SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction
- ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
- Much Ado About Noising: Dispelling the Myths of Generative Robotic Control
- GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
- DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
- TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation
- CycleManip: Enabling Cyclic Task Manipulation via Effective Historical Perception and Understanding
- MM-ACT: Learn from Multimodal Parallel Generation to Act
- Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning
- Image Generation as a Visual Planner for Robotic Manipulation
- LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
- Automated Generation of MDPs Using Logic Programming and LLMs for Robotic Applications
- Distracted Robot: How Visual Clutter Undermine Robotic Manipulation
- Improving Robotic Manipulation Robustness via NICE Scene Surgery
- DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
- LLM-Based Generalizable Hierarchical Task Planning and Execution for Heterogeneous Robot Teams with Event-Driven Replanning
- VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
- E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Continuized Discrete Diffusion
- From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
- Reinforcing Action Policies by Prophesying
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation
- AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
- Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation
- Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
- MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
- EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
- ArticFlow: Generative Simulation of Articulated Mechanisms
- Contact-Rich Robotic Assembly in Construction via Diffusion Policy Learning
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- DynaMimicGen: A Data Generation Framework for Robot Learning of Dynamic Tasks
- VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
- Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
- Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-the-Wild Human Demonstrations
- InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
- FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
- When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- Bi-AQUA: Bilateral Control-Based Imitation Learning for Underwater Robot Arms via Lighting-Aware Action Chunking with Transformers
- HMC: Learning Heterogeneous Meta-Control for Contact-Rich Loco-Manipulation
- Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
- AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
- RoboTidy : A 3D Gaussian Splatting Household Tidying Benchmark for Embodied Navigation and Action
- From Power to Precision: Learning Fine-grained Dexterity for Multi-fingered Robotic Hands
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
- Scalable Policy Evaluation with Video World Models
- Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
- Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
- Learning a Thousand Tasks in a Day
- Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
- SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
- RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
- Balance Equation-based Distributionally Robust Offline Imitation Learning
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
- ViPRA: Video Prediction for Robot Actions
- How Do VLAs Effectively Inherit from VLMs?
- 10 Open Challenges Steering the Future of Vision-Language-Action Models
- Lite VLA: Efficient Vision-Language-Action Control on CPU-Bound Edge Robots
- EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
- TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
- Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation
- Visual Spatial Tuning
- Temporal Action Selection for Action Chunking
- LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation
- Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
- Dexterous Robotic Piano Playing at Scale
- EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations
- A Step Toward World Models: A Survey on Robotic Manipulation
- BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
- Toward Accurate Long-Horizon Robotic Manipulation: Language-to-Action with Foundation Models via Scene Graphs
- Human-in-the-loop Online Rejection Sampling for Robotic Manipulation
- Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
- π_
RL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models - Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models
- Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
- When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models
- Enfold: Folding World-Generator Computation into Predictive Representations for Efficient Embodied Control
- Vision-TL-Action: Neuro-Symbolic Trajectory Generation from Visual Observations and Temporal Logic
- Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels
- PACE: Phase-Aware Chunk Execution for Robot Policies with Action Chunking
- Nautilus: From One Prompt to Plug-and-Play Robot Learning
- Integrative neurocybernetic modeling in the era of large-scale neuroscience
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
- DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI
- Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation
- Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World
- MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context
- ACG: Action Coherence Guidance for Flow-based VLA models
- Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
- PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- Semantic World Models
- VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting
- A Compositional Paradigm for Foundation Models: Towards Smarter Robotic Agents
- MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation
- RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation
- Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots
- Implicit State Estimation via Video Replanning
- From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- End-to-end Listen, Look, Speak and Act
- A Comprehensive Survey on World Models for Embodied AI
- RAPID Hand Prototype: Design of an Affordable, Fully-Actuated Biomimetic Hand for Dexterous Teleoperation
- VO-DP: Semantic-Geometric Adaptive Diffusion Policy for Vision-Only Robotic Manipulation
- QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
- Open TeleDex: A Hardware-Agnostic Teleoperation System for Imitation Learning based Dexterous Manipulation
- Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- ALOHA2 Robot Kitchen Application Scenario Reproduction Report
- RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
- Learning to Grasp Anything by Playing with Random Toys
- Reflection-Based Task Adaptation for Self-Improving VLA
- CoRA: Covariate-Aware Adaptation of Time Series Foundation Models
- Fast Visuomotor Policy for Robotic Manipulation
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- ManiAgent: An Agentic Framework for General Robotic Manipulation
- DemoHLM: From One Demonstration to Generalizable Humanoid Loco-Manipulation
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- TabVLA: Targeted Backdoor Attacks on Vision-Language-Action Models
- UniCoD: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
- ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- Dejavu: Towards Experience Feedback Learning for Embodied Intelligence
- Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
- VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
- NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos
- Maximum In-Support Return Modeling for Dynamic Recommendation with Language Model Prior
- IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- FLEET: Formal Language-Grounded Scheduling for Heterogeneous Robot Teams
- ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL Problems
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
- The Safety Challenge of World Models for Embodied AI Agents: A Review
- INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models
- VENTURA: Adapting Image Diffusion Models for Unified Task Conditioned Navigation
- MetaVLA: Unified Meta Co-training For Efficient Embodied Adaption
- FORGE-Tree: Diffusion-Forcing Tree Search for Long-Horizon Robot Manipulation
- BuilderBench -- A benchmark for generalist agents
- ARRC: Advanced Reasoning Robot Control - Knowledge-Driven Autonomous Manipulation Using Retrieval-Augmented Generation
- EmbodiedCoder: Parameterized Embodied Mobile Manipulation via Modern Coding Model
- D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
- Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
- SITCOM: Scaling Inference-Time COMpute for VLAs
- NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
- GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
- 3D-CovDiffusion: 3D-Aware Diffusion Policy for Coverage Path Planning
- Hybrid Training for Vision-Language-Action Models
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- RoboPilot: Generalizable Dynamic Robotic Manipulation with Dual-thinking Modes
- MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- Seeing Space and Motion: Enhancing Latent Actions with Spatial and Dynamic Awareness for VLA
- Reinforced Embodied Planning with Verifiable Reward for Real-World Robotic Manipulation
- Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
- Best of Sim and Real: Decoupled Visuomotor Manipulation via Learning Control in Simulation and Perception in Real
- Data-Efficient Multitask DAgger
- AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
- From Code to Action: Hierarchical Learning of Diffusion-VLM Policies
- FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
- Fidelity-Aware Data Composition for Robust Robot Generalization
- IA-VLA: Input Augmentation for Vision-Language-Action models in settings with semantically complex tasks
- Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
- GLUE: Global-Local Unified Encoding for Imitation Learning via Key-Patch Tracking
- Pixel Motion Diffusion is What We Need for Robot Control
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
- UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
- Developing Vision-Language-Action Model from Egocentric Videos
- DynaNav: Dynamic Feature and Layer Selection for Efficient Visual Navigation
- SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
- MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation
- RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models
- KeyWorld: Key Frame Reasoning Enables Effective and Efficient World Models
- Normalizing Flows are Capable Models for Bi-manual Visuomotor Policy
- CAD-Tokenizer: Towards Text-based CAD Prototyping via Modality-Specific Tokenization
- Large Pre-Trained Models for Bimanual Manipulation in 3D
- mindmap: Spatial Memory in Deep Feature Maps for 3D Action Policies
- One Filters All: A Generalist Filter for State Estimation
- Embodied AI: From LLMs to World Models
- Generalist Robot Manipulation beyond Action Labeled Data
- FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
- EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data
- RoboBRIDGE: A Modular Framework for Bridging Policies to Robust Real-World Robotic Agents
- Cross-Embodiment Transfer via Behavior-Aligned Representations
- LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation
- MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
- Self-evolved Imitation Learning in Simulated World
- SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation
- Do You Need Proprioceptive States in Visuomotor Policies?
- VGGT-DP: Generalizable Robot Control via Vision Foundation Models
- Residual Off-Policy RL for Finetuning Behavior Cloning Policies
- 3D Flow Diffusion Policy: Visuomotor Policy Learning via Generating Flow in 3D Space
- SEBVS: Synthetic Event-based Visual Servoing for Robot Navigation and Manipulation
- IDfRA: Self-Verification for Iterative Design in Robotic Assembly
- History-Aware Visuomotor Policy Learning via Point Tracking
- No Need for Real 3D: Fusing 2D Vision with Pseudo 3D Representations for Robotic Manipulation Learning
- TranTac: Leveraging Transient Tactile Signals for Contact-Rich Robotic Manipulation
- A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
- GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation
- Mental Accounts for Actions: EWA-Inspired Attention in Decision Transformers
- LodeStar: Long-horizon Dexterity via Synthetic Data Augmentation from Human Demonstrations
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- ExT: Towards Scalable Autonomous Excavation via Large-Scale Multi-Task Pretraining and Fine-Tuning
- CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human
- Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI
- VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation
- Toward Embodiment Equivariant Vision-Language-Action Policy
- Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models
- SeqVLA: Sequential Task Execution for Long-Horizon Manipulation with Completion-Aware Vision-Language-Action Model
- GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
- PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models
- Bridging Perception and Planning: Towards End-to-End Planning for Signal Temporal Logic Tasks
- Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
- Igniting VLMs toward the Embodied Space
- Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
- FEWT: Improving Humanoid Robot Perception with Frequency-Enhanced Wavelet-based Transformers
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
- Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models
- NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction
- Text2Touch: Tactile In-Hand Manipulation with LLM-Designed Reward Functions
- LLaDA-VLA: Vision Language Diffusion Action Models
- F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
- FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies
- COMMET: A System for Human-Induced Conflicts in Mobile Manipulation of Everyday Tasks
- EMMA: Scaling Mobile Manipulation via Egocentric Human Data
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
- FPC-VLA: A Vision-Language-Action Framework with a Supervisor for Failure Prediction and Correction
- ANNIE: Be Careful of Your Robots
- Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
- Learning to Ask: Decision Transformers for Adaptive Quantitative Group Testing
- Plantbot: Integrating Plant and Robot through LLM Modular Agent Networks
- MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation
- Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation
- EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control
- CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
- Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
- Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- An LLM-powered Natural-to-Robotic Language Translation Framework with Correctness Guarantees
- HyperTASR: Hypernetwork-Driven Task-Aware Scene Representations for Robust Manipulation
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Multimodal Data Storage and Retrieval for Embodied AI: A Survey
- Precise Action-to-Video Generation Through Visual Action Prompts
- Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Improving Pre-Trained Vision-Language-Action Policies with Model-Based Search
- Self-Guided Action Diffusion
- Human Centric General Physical Intelligence for Agile Manufacturing Automation
- TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
- Inspire or Predict? Exploring New Paradigms in Assisting Classical Planners with Large Language Models
- Agentic Design Review System
- KDPE: A Kernel Density Estimation Strategy for Diffusion Policy Trajectory Selection
- MLM: Learning Multi-task Loco-Manipulation Whole-Body Control for Quadruped Robot with Arm
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
- Leveraging OS-Level Primitives for Robotic Action Management
- Large Scale Robotic Material Handling: Learning, Planning, and Control
- DeepFleet: Multi-Agent Foundation Models for Mobile Robots
- GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
- ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks
- AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning
- SwarmVLM: VLM-Guided Impedance Control for Autonomous Navigation of Heterogeneous Robots in Dynamic Warehousing
- MolmoAct: Action Reasoning Models that can Reason in Space
- GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
- Multimodal learning with next-token prediction for large multimodal models
- Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
- PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation
- ASkDAgger: Active Skill-level Data Aggregation for Interactive Imitation Learning
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
- A tutorial note on collecting simulated data for vision-language-action models
- CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation
- ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
- ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
- CLASS: Contrastive Learning via Action Sequence Supervision for Robot Manipulation
- COLLAGE: Adaptive Fusion-based Retrieval for Augmented Policy Learning
- On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
- Video Generators are Robot Policies
- RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
- Policy Learning from Large Vision-Language Model Feedback without Reward Modeling
- From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
Related