OpenVLA: An Open-Source Vision-Language-Action Model
2024/06/13 by Moo Jin Kim, Kim, Moo Jin, Karl Pertsch +33 · 620 citations
Computer Science · #Semantic Web and Ontologies
paper · pdf · doi:10.48550/arxiv.2406.09246
Abstract
Large policies pretrained on a combination of Internet-scale vision-language data and diverse robot demonstrations have the potential to change how we teach robots new skills: rather than training new behaviors from scratch, we can fine-tune such vision-language-action (VLA) models to obtain robust, generalizable policies for visuomotor control. Yet, widespread adoption of VLAs for robotics has been challenging as 1) existing VLAs are largely closed and inaccessible to the public, and 2) prior work fails to explore methods for efficiently fine-tuning VLAs for new tasks, a key component for adoption. Addressing these challenges, we introduce OpenVLA, a 7B-parameter open-source VLA trained on a diverse collection of 970k real-world robot demonstrations. OpenVLA builds on a Llama 2 language model combined with a visual encoder that fuses pretrained features from DINOv2 and SigLIP. As a product of the added data diversity and new model components, OpenVLA demonstrates strong results for generalist manipulation, outperforming closed models such as RT-2-X (55B) by 16.5% in absolute task success rate across 29 tasks and multiple robot embodiments, with 7x fewer parameters. We further show that we can effectively fine-tune OpenVLA for new settings, with especially strong generalization results in multi-task environments involving multiple objects and strong language grounding abilities, and outperform expressive from-scratch imitation learning methods such as Diffusion Policy by 20.4%. We also explore compute efficiency; as a separate contribution, we show that OpenVLA can be fine-tuned on consumer GPUs via modern low-rank adaptation methods and served efficiently via quantization without a hit to downstream success rate. Finally, we release model checkpoints, fine-tuning notebooks, and our PyTorch codebase with built-in support for training VLAs at scale on Open X-Embodiment datasets.
Cited by
- τ: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision
- A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference
- From Generated Human Videos to Physically Plausible Robot Trajectories
- ProGuard: Towards Proactive Multimodal Safeguard
- Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
- VGGT-Ω
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- WorldDiT: A Unified Diffusion Architecture for World and Action Modeling
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
- Emergence of Human to Robot Transfer in Vision-Language-Action Models
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- GamiBench: Evaluating Spatial Reasoning and 2D-to-3D Planning Capabilities of MLLMs with Origami Folding Tasks
- MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model
- N0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
- N0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
- PAC-DP: PAC-Bayesian Diffusion Policy Learning
- LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratory
- DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
- Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
- GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
- CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model
- A Causality-aware Infer-diagnose-refine Framework for Test-time Modality Adaptation in VLA Models
- Teaching Tiny VLA Models Where to Look and How to Move
- RoboMME-Interference: Benchmarking Robot Memory Under Interference
- The Cartesian Cut in Agentic AI
- Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models
- Flexible Multitask Learning with Factorized Diffusion Policy
- StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
- RoboCade: Gamifying Robot Data Collection
- EVE: A Generator-Verifier System for Generative Policies
- LoLA: Long Horizon Latent Action Learning for General Robot Manipulation
- Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation
- Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting
- ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
- REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation
- MaP-AVR: A Meta-Action Planner for Agents Leveraging Vision Language Models and Retrieval-Augmented Generation
- Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface
- IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments
- Point What You Mean: Visually Grounded Instruction Policy
- STORM: Search-Guided Generative World Models for Robotic Manipulation
- AOMGen: Photoreal, Physics-Consistent Demonstration Generation for Articulated Object Manipulation
- Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
- SurgiPose: Estimating Surgical Tool Kinematics from Monocular Video for Surgical Robot Learning
- Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- Unifying Deep Predicate Invention with Pre-trained Foundation Models
- AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning
- Vidarc: Embodied Video Diffusion Model for Closed-loop Control
- Flying in Clutter on Monocular RGB by Learning in 3D Radiance Fields with Domain Adaptation
- Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
- ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning
- GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
- PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
- Spatia: Video Generation with Updatable Spatial Memory
- EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
- CaFe-TeleVision: A Coarse-to-Fine Teleoperation System with Immersive Situated Visualization for Enhanced Ergonomics
- DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
- BLURR: A Boosted Low-Resource Inference for Vision-Language-Action Models
- Sample-Efficient Robot Skill Learning for Construction Tasks: Benchmarking Hierarchical Reinforcement Learning and Vision-Language-Action VLA Model
- OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning
- Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
- Motus: A Unified Latent Action World Model
- SAGA: Open-World Mobile Manipulation via Structured Affordance Grounding
- Benchmarking the Generality of Vision-Language-Action Models
- Towards Logic-Aware Manipulation: A Knowledge Primitive for VLM-Based Assistants in Smart Manufacturing
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
- Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- Towards Accessible Physical AI: LoRA-Based Fine-Tuning of VLA Models for Real-World Robot Control
- Openpi Comet: Competition Solution For 2025 BEHAVIOR Challenge
- HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
- Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning
- GLaD: Geometric Latent Distillation for Vision-Language-Action Models
- Data-driven Interpretable Hybrid Robot Dynamics
- Language-Conditioned Safe Trajectory Generation for Spacecraft Rendezvous
- VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer
- Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging
- Bridging Scale Discrepancies in Robotic Control via Language-Based Action Representations
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- VLD: Visual Language Goal Distance for Reinforcement Learning Navigation
- Using Vision-Language Models as Proxies for Social Intelligence in Human-Robot Interaction
- See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
- Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulation
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- WAM-Diff: A Masked Diffusion VLA Framework with MoE and Online Reinforcement Learning for Autonomous Driving
- Training-Time Action Conditioning for Efficient Real-Time Chunking
- HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
- STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models
- FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
- MOVE: A Simple Motion-Based Data Collection Paradigm for Spatial Generalization in Robotic Manipulation
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- Embodied Co-Design for Rapidly Evolving Agents: Taxonomy, Frontiers, and Challenges
- Vision-Language-Action Models for Selective Robotic Disassembly: A Case Study on Critical Component Extraction from Desktops
- FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination
- SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL
- Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
- PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
- VAT: Vision Action Transformer by Unlocking Full Representation of ViT
- RoboScape-R: Unified Reward-Observation World Models for Generalizable Robotics Training via RL
- Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
- CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
- Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
- SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction
- ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
- Much Ado About Noising: Dispelling the Myths of Generative Robotic Control
- GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
- DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
- IGen: Scalable Data Generation for Robot Learning from Open-World Images
- VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
- CycleManip: Enabling Cyclic Task Manipulation via Effective Historical Perception and Understanding
- MM-ACT: Learn from Multimodal Parallel Generation to Act
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
- Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
- Sigma: The Key for Vision-Language-Action Models toward Telepathic Alignment
- RealAppliance: Let High-fidelity Appliance Assets Controllable and Workable as Aligned Real Manuals
- LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
- SafeHumanoid: VLM-RAG-driven Control of Upper Body Impedance for Humanoid Robot
- Distracted Robot: How Visual Clutter Undermine Robotic Manipulation
- Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
- BINDER: Instantly Adaptive Mobile Manipulation with Open-Vocabulary Commands
- TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos
- DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
- LLM-Based Generalizable Hierarchical Task Planning and Execution for Heterogeneous Robot Teams with Event-Driven Replanning
- Attention-Guided Patch-Wise Sparse Adversarial Attacks on Vision-Language-Action Models
- DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
- VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
- E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Continuized Discrete Diffusion
- From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings
- Hyper-GoalNet: Goal-Conditioned Manipulation Policy Learning with HyperNetworks
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
- Dynamic Test-Time Compute Scaling in Control Policy: Difficulty-Aware Stochastic Interpolant Policy
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Reinforcing Action Policies by Prophesying
- Arcadia: Toward a Full-Lifecycle Framework for Embodied Lifelong Learning
- Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation
- DeeAD: Dynamic Early Exit of Vision-Language Action for Efficient Autonomous Driving
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
- Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation
- Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
- MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
- Mixture of Horizons in Action Chunking
- EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
- ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
- MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation
- IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
- VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
- H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation
- Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
- Unify Robot Actions in Camera Frame
- METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
- SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
- InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
- FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
- When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- In-N-On: Scaling Egocentric Manipulation with in-the-wild and on-task Data
- SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
- Theoretical Closed-loop Stability Bounds for Dynamical System Coupled with Diffusion Policies
- Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception
- HMC: Learning Heterogeneous Meta-Control for Contact-Rich Loco-Manipulation
- AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
- AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models
- FoldPath: End-to-End Object-Centric Motion Generation via Modulated Implicit Paths
- Decoupled Action Head: Confining Task Knowledge to Conditioning Layers
- VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
- AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
- Scalable Policy Evaluation with Video World Models
- Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
- Phantom Menace: Exploring and Enhancing the Robustness of VLA Models Against Physical Sensor Attacks
- AdaptPNP: Integrating Prehensile and Non-Prehensile Skills for Adaptive Robotic Manipulation
- SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
- RoboBenchMart: Benchmarking Robots in Retail Environment
- Learning a Thousand Tasks in a Day
- Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
- SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
- SPIDER: Scalable Physics-Informed Dexterous Retargeting
- RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation
- RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
- DRACO: Co-design for DSP-Efficient Rigid Body Dynamics Accelerator
- ViPRA: Video Prediction for Robot Actions
- SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
- How Do VLAs Effectively Inherit from VLMs?
- ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
- Towards Human-AI-Robot Collaboration and AI-Agent based Digital Twins for Parkinson's Disease Management: Review and Outlook
- From Words to Safety: Language-Conditioned Safety Filtering for Robot Navigation
- EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
- TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
- Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation
- Visual Spatial Tuning
- Cambrian-S: Towards Spatial Supersensing in Video
- Real-to-Sim Robot Policy Evaluation with Gaussian Splatting Simulation of Soft-Body Interactions
- Embodiment Transfer Learning for Vision-Language-Action Models
- GraSP-VLA: Graph-based Symbolic Action Representation for Long-Horizon Planning with VLA Policies
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- GUIDES: Guidance Using Instructor-Distilled Embeddings for Pre-trained Robot Policy Enhancement
- Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
- LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation
- Learning Interactive World Model for Object-Centric Reinforcement Learning
- OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
- iFlyBot-VLA Technical Report
- EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations
- Towards a Multi-Embodied Grasping Agent
- Learning Generalizable Visuomotor Policy through Dynamics-Alignment
- A Step Toward World Models: A Survey on Robotic Manipulation
- End-to-End Dexterous Arm-Hand VLA Policies via Shared Autonomy: VR Teleoperation Augmented by Autonomous Hand VLA Policy for Efficient Data Collection
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
- Hybrid Consistency Policy: Decoupling Multi-Modal Diversity and Real-Time Efficiency in Robotic Manipulation
- RoboOS-NeXT: A Unified Memory-based Framework for Lifelong, Scalable, and Robust Multi-Robot Collaboration
- Human-in-the-loop Online Rejection Sampling for Robotic Manipulation
- Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
- π_
RL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models - Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models
- When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning
- SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models
- Enfold: Folding World-Generator Computation into Predictive Representations for Efficient Embodied Control
- Vision-TL-Action: Neuro-Symbolic Trajectory Generation from Visual Observations and Temporal Logic
- StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
- Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels
- The missing data for intelligent scientific instruments
- Route by Kinematics, Act by Observation: Kinematics-Supervised Expert Routing in MoE-Augmented VLA
- Practice Makes Policies: Bootstrapping and Consolidating Robotic Capabilities from Zero Human Demonstrations
- MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
- ZeroShotOpt: Towards Zero-Shot Pretrained Models for Efficient Black-Box Optimization
- Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method
- VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision
- PACE: Phase-Aware Chunk Execution for Robot Policies with Action Chunking
- InDex: Empowering VLA Models with Intent-Conditioned Arm-Hand Coordination for Dexterous Manipulation
- Can Vision-Language-Action Models Learn from Real-World Data Continually without Forgetting?
- Expanding LLM Agent Boundaries with Strategy-Guided Exploration
- TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans
- Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals via Vision-Language Models
- CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs
- Blindfolded Experts Generalize Better: Insights from Robotic Manipulation and Videogames
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI
- Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context
- A Survey on Efficient Vision-Language-Action Models
- UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- Dexbotic: Open-Source Vision-Language-Action Toolbox
- Reliable Robotic Task Execution in the Face of Anomalies
- OmniDexGrasp: Generalizable Dexterous Grasping via Foundation Model and Force Feedback
- ACG: Action Coherence Guidance for Flow-based VLA models
- Generalizable Hierarchical Skill Learning via Object-Centric Representation
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- Exploring Conditions for Diffusion models in Robotic Control
- Learning Affordances at Inference-Time for Vision-Language-Action Models
- Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
- GigaBrain-0: A World Model-Powered Vision-Language-Action Model
- Semantic World Models
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting
- A Compositional Paradigm for Foundation Models: Towards Smarter Robotic Agents
- EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and Retrieval
- MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning
- RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation
- From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
- RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- Learning to play: A Multimodal Agent for 3D Game-Play
- End-to-end Listen, Look, Speak and Act
- RM-RL: Role-Model Reinforcement Learning for Precise Robot Manipulation
- DexCanvas: Bridging Human Demonstrations and Robot Learning for Dexterous Manipulation
- RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon Tasks
- VLA2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
- GOPLA: Generalizable Object Placement Learning via Synthetic Augmentation of Human Arrangement
- Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
- DeDelayed: Deleting Remote Inference Delay via On-Device Correction
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
- RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation
- Model-agnostic Adversarial Attack and Defense for Vision-Language-Action Models
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
- From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails
- Reasoning in Space via Grounding in the World
- VLA-0: Building State-of-the-Art VLAs with Zero Modification
- Learning to Grasp Anything by Playing with Random Toys
- Reflection-Based Task Adaptation for Self-Improving VLA
- Fast Visuomotor Policy for Robotic Manipulation
- Improving Generative Behavior Cloning via Self-Guidance and Adaptive Chunking
- EmboMatrix: A Scalable Training-Ground for Embodied Decision-Making
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- ARMADA: Autonomous Online Failure Detection and Human Shared Control Empower Scalable Real-world Deployment and Adaptation
- ManiAgent: An Agentic Framework for General Robotic Manipulation
- HiMaCon: Discovering Hierarchical Manipulation Concepts from Unlabeled Multi-Modal Data
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- A Survey on Agentic Multimodal Large Language Models
- In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
- TabVLA: Targeted Backdoor Attacks on Vision-Language-Action Models
- More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
- UniCoD: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
- High-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning with Gaussian Splatting
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- Dejavu: Towards Experience Feedback Learning for Embodied Intelligence
- Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- An agentic artificially intelligent X-ray scientist
- VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
- Failure Prediction at Runtime for Generative Robot Policies
- Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
- Placeit! A Framework for Learning Robot Object Placement Skills
- Goal-oriented Backdoor Attack against Vision-Language-Action Models via Physical Objects
- PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
- Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos
- R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation
- USIM and U0: A Vision-Language-Action Dataset and Model for General Underwater Robots
- Executable Analytic Concepts as the Missing Link Between VLM Insight and Precise Manipulation
- IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
- BLAZER: Bootstrapping LLM-based Manipulation Agents with Zero-Shot Data Generation
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL Problems
- TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
- RLinf-VLA: A Unified and Efficient Framework for VLA+RL Training
- VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
- Medical Vision Language Models as Policies for Robotic Surgery
- The Safety Challenge of World Models for Embodied AI Agents: A Review
- INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models
- MetaVLA: Unified Meta Co-training For Efficient Embodied Adaption
- FORGE-Tree: Diffusion-Forcing Tree Search for Long-Horizon Robot Manipulation
- D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
- Verifier-free Test-Time Sampling for Vision-Language-Action Models
- Active Semantic Perception
- Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
- Exploring OCR-augmented Generation for Bilingual VQA
- Activation Quantization of Vision Encoders Needs Prefixing Registers
- Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- Contrastive Representation Regularization for Vision-Language-Action Models
- VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
- ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
- GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
- HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
- Hybrid Training for Vision-Language-Action Models
- Traj2Action: A Co-Denoising Framework for Trajectory-Guided Human-to-Robot Skill Transfer
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
- Seeing Space and Motion: Enhancing Latent Actions with Spatial and Dynamic Awareness for VLA
- Accelerating Transformers in Online RL
- Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive Policies
- Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
- VLA Model Post-Training via Action-Chunked PPO and Self Behavior Cloning
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation
- dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
- Data-Efficient Multitask DAgger
- Scaling Synthetic Task Generation for Agents via Exploration
- From Code to Action: Hierarchical Learning of Diffusion-VLM Policies
- Preference-Based Long-Horizon Robotic Stacking with Multimodal Large Language Models
- World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training
- Emergent World Representations in OpenVLA
- Fidelity-Aware Data Composition for Robust Robot Generalization
- PhysiAgent: An Embodied Agent Framework in Physical World
- IA-VLA: Input Augmentation for Vision-Language-Action models in settings with semantically complex tasks
- Training Agents Inside of Scalable World Models
- AutoPrune: Each Complexity Deserves a Pruning Policy
- Control Your Robot: A Unified System for Robot Control and Policy Deployment
- Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
- Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints
- OVSeg3R: Learn Open-vocabulary Instance Segmentation from 2D via 3D Reconstruction
- Space Robotics Bench: Robot Learning Beyond Earth
- GLUE: Global-Local Unified Encoding for Imitation Learning via Key-Patch Tracking
- LAGEA: Language Guided Embodied Agents for Robotic Manipulation
- Transferring Vision-Language-Action Models to Industry Applications: Architectures, Performance, and Challenges
- MMPB: It's Time for Multi-Modal Personalization
- Pixel Motion Diffusion is What We Need for Robot Control
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
- WoW: Towards a World omniscient World model Through Embodied Interaction
- IROSA: Interactive Robot Skill Adaptation Using Natural Language
- ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation
- RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
- From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment
- Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
- Developing Vision-Language-Action Model from Egocentric Videos
- SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
- MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation
- VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
- RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models
- AnywhereVLA: Language-Conditioned Exploration and Mobile Manipulation
- ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation
- mindmap: Spatial Memory in Deep Feature Maps for 3D Action Policies
- Parse-Augment-Distill: Learning Generalizable Bimanual Visuomotor Policies from Single Human Video
- One Filters All: A Generalist Filter for State Estimation
- Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
- Embodied AI: From LLMs to World Models
- Generalist Robot Manipulation beyond Action Labeled Data
- RoboSSM: Scalable In-context Imitation Learning via State-Space Models
- Beyond Human Demonstrations: Diffusion-Based Reinforcement Learning to Generate Data for VLA Training
- FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
- Diffusion-Based Impedance Learning for Contact-Rich Manipulation Tasks
- EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data
- Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
- RoboBRIDGE: A Modular Framework for Bridging Policies to Robust Real-World Robotic Agents
- SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
- World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
- Cross-Embodiment Transfer via Behavior-Aligned Representations
- RL2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy
- System 0/1/2/3: Quad-Process Theory for Multitimescale Embodied Collective Cognitive Systems
- SEBVS: Synthetic Event-based Visual Servoing for Robot Navigation and Manipulation
- In vivo feasibility study of humanoid robots in surgery
- HLG: Comprehensive 3D Room Construction via Hierarchical Layout Generation
- OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation
- SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
- Position: Human-Robot Interaction in Embodied Intelligence Demands a Shift From Static Privacy Controls to Dynamic Learning
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- Eva-VLA: Evaluating Vision-Language-Action Models' Robustness Under Real-World Physical Variations
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- N2M: Bridging Navigation and Manipulation by Learning Pose Preference from Rollout
- Do You Need Proprioceptive States in Visuomotor Policies?
- Growing with Your Embodied Agent: A Human-in-the-Loop Lifelong Code Generation Framework for Long-Horizon Manipulation Skills
- VGGT-DP: Generalizable Robot Control via Vision Foundation Models
- Residual Off-Policy RL for Finetuning Behavior Cloning Policies
- SINGER: An Onboard Generalist Vision-Language Navigation Policy for Drones
- 3D Flow Diffusion Policy: Visuomotor Policy Learning via Generating Flow in 3D Space
- Query-Centric Diffusion Policy for Generalizable Robotic Assembly
- Latent Action Pretraining Through World Modeling
- PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation
- DINOv3-Diffusion Policy: Self-Supervised Large Visual Model for Visuomotor Diffusion Policy Learning
- Prepare Before You Act: Learning From Humans to Rearrange Initial States
- RoboSeek: You Need to Interact with Your Objects
- I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
- History-Aware Visuomotor Policy Learning via Point Tracking
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- No Need for Real 3D: Fusing 2D Vision with Pseudo 3D Representations for Robotic Manipulation Learning
- KV-Efficient VLA: A Method to Speed up Vision Language Models with RNN-Gated Chunked KV Cache
- TranTac: Leveraging Transient Tactile Signals for Contact-Rich Robotic Manipulation
- Video-to-BT: Generating Reactive Behavior Trees from Human Demonstration Videos for Robotic Assembly
- Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
- Factorizing Diffusion Policies for Observation Modality Prioritization
- Dynamic Objects Relocalization in Changing Environments with Flow Matching
- See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
- CoReVLA: A Dual-Stage End-to-End Autonomous Driving Framework for Long-Tail Scenarios via Collect-and-Refine
- A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
- GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation
- Compose by Focus: Scene Graph-based Atomic Skills
- RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- How Good are Foundation Models in Step-by-Step Embodied Reasoning?
- Self-Improving Embodied Foundation Models
- Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue
- ExT: Towards Scalable Autonomous Excavation via Large-Scale Multi-Task Pretraining and Fine-Tuning
- Robot Control Stack: A Lean Ecosystem for Robot Learning at Scale
- CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human
- Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- COMPASS: Confined-space Manipulation Planning with Active Sensing Strategy
- RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI
- VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation
- Toward Embodiment Equivariant Vision-Language-Action Policy
- DreamControl: Human-Inspired Whole-Body Humanoid Control for Scene Interaction via Guided Diffusion
- CLAW: A Vision-Language-Action Framework for Weight-Aware Robotic Grasping
- Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models
- Dual-Actor Fine-Tuning of VLA Models: A Talk-and-Tweak Human-in-the-Loop Approach
- Language Conditioning Improves Accuracy of Aircraft Goal Prediction in Non-Towered Airspace
- SeqVLA: Sequential Task Execution for Long-Horizon Manipulation with Completion-Aware Vision-Language-Action Model
- GestOS: Advanced Hand Gesture Interpretation via Large Language Models to control Any Type of Robot
- GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
- PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models
- Robust Online Residual Refinement via Koopman-Guided Dynamics Modeling
- HARMONIC: A Content-Centric Cognitive Robotic Architecture
- From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
- The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
- Bridging Perception and Planning: Towards End-to-End Planning for Signal Temporal Logic Tasks
- Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
- Embodied Navigation Foundation Model
- EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models
- Igniting VLMs toward the Embodied Space
- AssemMate: Graph-Based LLM for Robotic Assembly Assistance
- Cross-Platform Scaling of Vision-Language-Action Models from Edge to Cloud GPUs
- Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
- An integrated process for design and control of lunar robotics using AI and simulation
- ActivePose: Active 6D Object Pose Estimation and Tracking for Robotic Manipulation
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
- Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- Self-Augmented Robot Trajectory: Efficient Imitation Learning via Safe Self-augmentation with Demonstrator-annotated Precision
- NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows
- TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making
- RoboMatch: A Unified Mobile-Manipulation Teleoperation Platform with Auto-Matching Network Architecture for Long-Horizon Tasks
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning
- TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
- RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction
- Graph-Fused Vision-Language-Action for Policy Reasoning in Multi-Arm Robotic Manipulation
- CRISP -- Compliant ROS2 Controllers for Learning-Based Manipulation Policies and Teleoperation
- LLaDA-VLA: Vision Language Diffusion Action Models
- CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis
- Robotic Manipulation Framework Based on Semantic Keypoints for Packing Shoes of Different Sizes, Shapes, and Softness
- GELATO: Multi-Instruction Trajectory Reshaping via Geometry-Aware Multiagent-based Orchestration
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
- FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies
- OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation
- COMMET: A System for Human-Induced Conflicts in Mobile Manipulation of Everyday Tasks
- EMMA: Scaling Mobile Manipulation via Egocentric Human Data
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
- FPC-VLA: A Vision-Language-Action Framework with a Supervisor for Failure Prediction and Correction
- Long-Horizon Visual Imitation Learning via Plan and Code Reflection
- ANNIE: Be Careful of Your Robots
- Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
- U-ARM : Ultra low-cost general teleoperation interface for robot manipulation
- OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- Do What? Teaching Vision-Language-Action Models to Reject the Impossible
- Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
- Articulated Object Estimation in the Wild
- MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation
- Galaxea Open-World Dataset and G0 Dual-System VLA Model
- Mechanistic interpretability for steering vision-language-action models
- ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
- Prompt-to-Product: Generative Assembly via Bimanual Manipulation
- EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control
- CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
- Embodied AI: Emerging Risks and Opportunities for Policy Action
- Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
- Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
- Ego-centric Predictive Model Conditioned on Hand Trajectories
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- HyperTASR: Hypernetwork-Driven Task-Aware Scene Representations for Robust Manipulation
- Survey of Vision-Language-Action Models for Embodied Manipulation
- TransLLM: A Unified Multi-Task Foundation Framework for Urban Transportation via Learnable Prompting
- The Social Context of Human-Robot Interactions
- CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- Sim-to-Real Dynamic Object Manipulation on Conveyor Systems via Optimization Path Shaping
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Improving Pre-Trained Vision-Language-Action Policies with Model-Based Search
- Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping
- OmniD: Generalizable Robot Manipulation Policy via Image-Based BEV Representation
- Human Centric General Physical Intelligence for Agile Manufacturing Automation
- ImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous Driving
- TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
- Multi-Group Equivariant Augmentation for Reinforcement Learning in Robot Manipulation
- KDPE: A Kernel Density Estimation Strategy for Diffusion Policy Trajectory Selection
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
- Leveraging OS-Level Primitives for Robotic Action Management
- GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation
- Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
- Masquerade: Learning from In-the-wild Human Videos using Data-Editing
- How Does a Virtual Agent Decide Where to Look? -- Symbolic Cognitive Reasoning for Embodied Head Rotation
- OmniVTLA: Vision-Tactile-Language-Action Model with Semantic-Aligned Tactile Sensing
- Boosting Action-Information via a Variational Bottleneck on Unlabelled Robot Videos
- GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
- Spatial Traces: Enhancing VLA Models with Spatial-Temporal Understanding
- ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks
- PCHands: PCA-based Hand Pose Synergy Representation on Manipulators with N-DoF
- DETACH: Cross-domain Learning for Long-Horizon Tasks via Mixture of Disentangled Experts
- MolmoAct: Action Reasoning Models that can Reason in Space
- AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies
- GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
- Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
- Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution
- Towards Balanced Behavior Cloning from Imbalanced Datasets
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- Information-Theoretic Graph Fusion with Vision-Language-Action Model for Policy Reasoning and Dual Robotic Control
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
- INTENTION: Inferring Tendencies of Humanoid Robot Motion Through Interactive Intuition and Grounded VLM
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
- Learning Robust Intervention Representations with Delta Embeddings
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- Static and Plugged: Make Embodied Evaluation Simple
- A tutorial note on collecting simulated data for vision-language-action models
- CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation
- Improving Generalization of Language-Conditioned Robot Manipulation
- FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models
- RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
- VLH: Vision-Language-Haptics Foundation Model
- On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
- HannesImitation: Grasping with the Hannes Prosthetic Hand via Imitation Learning
- XRoboToolkit: A Cross-Platform Framework for Robot Teleoperation
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
- Policy Learning from Large Vision-Language Model Feedback without Reward Modeling
- H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
- Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
- From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
Related