RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
2023/07/28 by Anthony Brohan, Noah Brown, Brohan, Anthony +105 · 548 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Robotics (cs.RO) #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2307.15818
openalex publication_date 2023/07/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning. Our goal is to enable a single end-to-end trained model to both learn to map robot observations to actions and enjoy the benefits of large-scale pretraining on language and vision-language data from the web. To this end, we propose to co-fine-tune state-of-the-art vision-language models on both robotic trajectory data and Internet-scale vision-language tasks, such as visual question answering. In contrast to other approaches, we propose a simple, general recipe to achieve this goal: in order to fit both natural language responses and robotic actions into the same format, we express the actions as text tokens and incorporate them directly into the training set of the model in the same way as natural language tokens. We refer to such category of models as vision-language-action models (VLA) and instantiate an example of such a model, which we call RT-2. Our extensive evaluation (6k evaluation trials) shows that our approach leads to performant robotic policies and enables RT-2 to obtain a range of emergent capabilities from Internet-scale training. This includes significantly improved generalization to novel objects, the ability to interpret commands not present in the robot training data (such as placing an object onto a particular number or icon), and the ability to perform rudimentary reasoning in response to user commands (such as picking up the smallest or largest object, or the one closest to another object). We further show that incorporating chain of thought reasoning allows RT-2 to perform multi-stage semantic reasoning, for example figuring out which object to pick up for use as an improvised hammer (a rock), or which type of drink is best suited for someone who is tired (an energy drink).
Cited by
- τ: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision
- A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference
- From Generated Human Videos to Physically Plausible Robot Trajectories
- VGGT-Ω
- Agentic Physical AI toward a Domain-Specific Foundation Model for Energy Systems: A Case Study on Nuclear Reactor Control
- Embodied Learning of Reward for Musculoskeletal Control with Vision Language Models
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- Emergence of Human to Robot Transfer in Vision-Language-Action Models
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model
- N0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
- The Semantic Least-Energy Principle: A Hypothesis for Intelligence
- LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratory
- Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems
- WCM: World-Cognition Model for Generalizable Human-Robot Interaction
- DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
- The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation
- Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
- A Causality-aware Infer-diagnose-refine Framework for Test-time Modality Adaptation in VLA Models
- Teaching Tiny VLA Models Where to Look and How to Move
- An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia
- Compositional Motion Generation from Demonstration with Object-Centric Neural Fields
- RoboMME-Interference: Benchmarking Robot Memory Under Interference
- PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments
- The Cartesian Cut in Agentic AI
- Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models
- StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
- AstraNav-World: World Model for Foresight Control and Consistency
- RoboCade: Gamifying Robot Data Collection
- LoLA: Long Horizon Latent Action Learning for General Robot Manipulation
- Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting
- ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
- Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation
- REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation
- MaP-AVR: A Meta-Action Planner for Agents Leveraging Vision Language Models and Retrieval-Augmented Generation
- Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface
- DeliveryBench: Can Agents Earn Profit in Real World?
- Point What You Mean: Visually Grounded Instruction Policy
- STORM: Search-Guided Generative World Models for Robotic Manipulation
- Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
- Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- Name That Part: 3D Part Segmentation and Naming
- Tiny Recursive Control: Iterative Reasoning for Efficient Optimal Control
- GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
- PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence
- VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- Large Video Planner Enables Generalizable Robot Control
- MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
- VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
- Spatia: Video Generation with Updatable Spatial Memory
- JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- Sample-Efficient Robot Skill Learning for Construction Tasks: Benchmarking Hierarchical Reinforcement Learning and Vision-Language-Action VLA Model
- OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning
- Multi-Robot Motion Planning from Vision and Language using Heat-Inspired Diffusion
- Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Motus: A Unified Latent Action World Model
- SAGA: Open-World Mobile Manipulation via Structured Affordance Grounding
- TS-DP: Reinforcement Speculative Decoding For Temporal Adaptive Diffusion Policy Acceleration
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- RoboNeuron: A Middle-Layer Infrastructure for Agent-Driven Orchestration in Embodied AI
- Openpi Comet: Competition Solution For 2025 BEHAVIOR Challenge
- Evaluating Gemini Robotics Policies in a Veo World Simulator
- HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
- Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning
- GLaD: Geometric Latent Distillation for Vision-Language-Action Models
- Fed-SE: Federated Self-Evolution for Privacy-Constrained Multi-Environment LLM Agents
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer
- Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging
- Dora: QoE-Aware Hybrid Parallelism for Distributed Edge AI
- Bridging Scale Discrepancies in Robotic Control via Language-Based Action Representations
- Semantic-Metric Bayesian Risk Fields: Learning Robot Safety from Human Videos with a VLM Prior
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
- Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulation
- Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge
- Invariance Co-training for Robot Visual Generalization
- Stitch and Tell: A Structured Multimodal Data Augmentation Method for Spatial Understanding
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models
- Correspondence-Oriented Imitation Learning: Flexible Visuomotor Control with 3D Conditioning
- HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
- STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models
- FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
- MOVE: A Simple Motion-Based Data Collection Paradigm for Spatial Generalization in Robotic Manipulation
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- LAWS: Learning from Actual Workloads Symbolically -- A Self-Certifying Parametrized Cache Architecture for Neural Inference, Robotics, and Edge Deployment
- Embodied Co-Design for Rapidly Evolving Agents: Taxonomy, Frontiers, and Challenges
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- dVLM-AD: Enhance Diffusion Vision-Language-Model for Driving via Controllable Reasoning
- Vision-Language-Action Models for Selective Robotic Disassembly: A Case Study on Critical Component Extraction from Desktops
- FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination
- BiTAgent: A Task-Aware Modular Framework for Bidirectional Coupling between Multimodal Large Language Models and World Models
- ResponsibleRobotBench: Benchmarking Responsible Robot Manipulation using Multi-modal Large Language Models
- Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
- PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
- VAT: Vision Action Transformer by Unlocking Full Representation of ViT
- PerFACT: Motion Policy with LLM-Powered Dataset Synthesis and Fusion Action-Chunking Transformers
- Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
- CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
- VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling
- Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
- SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction
- Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
- ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
- Much Ado About Noising: Dispelling the Myths of Generative Robotic Control
- GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
- DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
- Modality-Augmented Fine-Tuning of Foundation Robot Policies for Cross-Embodiment Manipulation on GR1 and G1
- VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
- CycleManip: Enabling Cyclic Task Manipulation via Effective Historical Perception and Understanding
- MM-ACT: Learn from Multimodal Parallel Generation to Act
- Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
- Sigma: The Key for Vision-Language-Action Models toward Telepathic Alignment
- REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories
- CC-FMO: Camera-Conditioned Zero-Shot Single Image to 3D Scene Generation with Foundation Model Orchestration
- SimScale: Learning to Drive via Real-World Simulation at Scale
- LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
- AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
- SafeHumanoid: VLM-RAG-driven Control of Upper Body Impedance for Humanoid Robot
- Adapting Like Humans: A Metacognitive Agent with Test-time Reasoning
- Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
- BINDER: Instantly Adaptive Mobile Manipulation with Open-Vocabulary Commands
- TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos
- LLM-Based Generalizable Hierarchical Task Planning and Execution for Heterogeneous Robot Teams with Event-Driven Replanning
- DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
- VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
- E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Continuized Discrete Diffusion
- From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
- Dynamic Test-Time Compute Scaling in Control Policy: Difficulty-Aware Stochastic Interpolant Policy
- ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation
- Arcadia: Toward a Full-Lifecycle Framework for Embodied Lifelong Learning
- Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
- MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action Generalization
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation
- Learning Massively Multitask World Models for Continuous Control
- AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
- Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation
- Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
- MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
- AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
- EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
- ArticFlow: Generative Simulation of Articulated Mechanisms
- MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- DynaMimicGen: A Data Generation Framework for Robot Learning of Dynamic Tasks
- VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
- Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
- Unify Robot Actions in Camera Frame
- PathAgent: Toward Interpretable Analysis of Whole-slide Pathology Images via Large Language Model-based Agentic Reasoning
- METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
- SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
- InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
- Bridging VLMs and Embodied Intelligence with Deliberate Practice Policy Optimization
- When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- Semantic Glitch: Agency and Artistry in an Autonomous Pixel Cloud
- SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
- IPR-1: Interactive Physical Reasoner
- HMC: Learning Heterogeneous Meta-Control for Contact-Rich Loco-Manipulation
- Enhancing End-to-End Autonomous Driving with Risk Semantic Distillaion from VLM
- AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
- An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
- AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models
- Decoupled Action Head: Confining Task Knowledge to Conditioning Layers
- Utilizing LLMs for Industrial Process Automation: A Case Study on Modifying RAPID Programs
- Phantom Menace: Exploring and Enhancing the Robustness of VLA Models Against Physical Sensor Attacks
- SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
- Learning a Thousand Tasks in a Day
- Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
- SPIDER: Scalable Physics-Informed Dexterous Retargeting
- RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation
- Hey Pentti, We Did (More of) It!: A Vector-Symbolic Lisp With Residue Arithmetic
- RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
- Balance Equation-based Distributionally Robust Offline Imitation Learning
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
- ViPRA: Video Prediction for Robot Actions
- How Do VLAs Effectively Inherit from VLMs?
- ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
- 10 Open Challenges Steering the Future of Vision-Language-Action Models
- Towards Human-AI-Robot Collaboration and AI-Agent based Digital Twins for Parkinson's Disease Management: Review and Outlook
- From Words to Safety: Language-Conditioned Safety Filtering for Robot Navigation
- VLM-driven Skill Selection for Robotic Assembly Tasks
- Lite VLA: Efficient Vision-Language-Action Control on CPU-Bound Edge Robots
- EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
- TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
- Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation
- Visual Spatial Tuning
- Real-to-Sim Robot Policy Evaluation with Gaussian Splatting Simulation of Soft-Body Interactions
- Embodiment Transfer Learning for Vision-Language-Action Models
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- ROSBag MCP Server: Analyzing Robot Data with LLMs for Agentic Embodied AI Applications
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
- LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation
- Dexterous Robotic Piano Playing at Scale
- iFlyBot-VLA Technical Report
- Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
- EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations
- A Step Toward World Models: A Survey on Robotic Manipulation
- End-to-End Dexterous Arm-Hand VLA Policies via Shared Autonomy: VR Teleoperation Augmented by Autonomous Hand VLA Policy for Efficient Data Collection
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
- BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
- Toward Accurate Long-Horizon Robotic Manipulation: Language-to-Action with Foundation Models via Scene Graphs
- EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities
- DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
- RoboOS-NeXT: A Unified Memory-based Framework for Lifelong, Scalable, and Robust Multi-Robot Collaboration
- Human-in-the-loop Online Rejection Sampling for Robotic Manipulation
- MossNet: Mixture of State-Space Experts is a Multi-Head Attention
- Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
- Pelican-VL 1.0: A Foundation Brain Model for Embodied Intelligence
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
- When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning
- SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models
- CG-World: A Large-Scale World-State Dataset and Protocol for World Models
- Enfold: Folding World-Generator Computation into Predictive Representations for Efficient Embodied Control
- Vision-TL-Action: Neuro-Symbolic Trajectory Generation from Visual Observations and Temporal Logic
- Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation
- Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels
- When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making
- Practice Makes Policies: Bootstrapping and Consolidating Robotic Capabilities from Zero Human Demonstrations
- PACE: Phase-Aware Chunk Execution for Robot Policies with Action Chunking
- InDex: Empowering VLA Models with Intent-Conditioned Arm-Hand Coordination for Dexterous Manipulation
- Nautilus: From One Prompt to Plug-and-Play Robot Learning
- Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
- Integrative neurocybernetic modeling in the era of large-scale neuroscience
- TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans
- SG-CoT: An Ambiguity-Aware Robotic Planning Framework using Scene Graph Representations
- Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals via Vision-Language Models
- What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- FruitProm: Probabilistic Maturity Estimation and Detection of Fruits and Vegetables
- SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space
- DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI
- Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation
- Human Machine Social Hybrid Intelligence:A Collaborative Decision Making Framework for Large Model Agent Groups and Human Experts
- Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
- A Survey on Efficient Vision-Language-Action Models
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation
- NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
- Dexbotic: Open-Source Vision-Language-Action Toolbox
- ACG: Action Coherence Guidance for Flow-based VLA models
- Towards Reliable Code-as-Policies: A Neuro-Symbolic Framework for Embodied Task Planning
- Out-of-Distribution Detection for Safety Assurance of AI and Autonomous Systems
- Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- MemER: Scaling Up Memory for Robot Control via Experience Retrieval
- DAIL: Beyond Task Ambiguity for Language-Conditioned Reinforcement Learning
- Learning Affordances at Inference-Time for Vision-Language-Action Models
- Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- A Compositional Paradigm for Foundation Models: Towards Smarter Robotic Agents
- EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and Retrieval
- Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models
- MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning
- RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation
- Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots
- From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
- RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- Learning to play: A Multimodal Agent for 3D Game-Play
- T3 Planner: A Self-Correcting LLM Framework for Robotic Motion Planning with Temporal Logic
- End-to-end Listen, Look, Speak and Act
- RAPID Hand Prototype: Design of an Affordable, Fully-Actuated Biomimetic Hand for Dexterous Teleoperation
- RM-RL: Role-Model Reinforcement Learning for Precise Robot Manipulation
- DexCanvas: Bridging Human Demonstrations and Robot Learning for Dexterous Manipulation
- RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon Tasks
- VLA2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
- QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
- Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
- InfraGPT Smart Infrastructure: An End-to-End VLM-Based Framework for Detecting and Managing Urban Defects
- DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
- ALOHA2 Robot Kitchen Application Scenario Reproduction Report
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
- VLA-0: Building State-of-the-Art VLAs with Zero Modification
- Learning to Grasp Anything by Playing with Random Toys
- Reflection-Based Task Adaptation for Self-Improving VLA
- Fast Visuomotor Policy for Robotic Manipulation
- EmboMatrix: A Scalable Training-Ground for Embodied Decision-Making
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- ManiAgent: An Agentic Framework for General Robotic Manipulation
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
- TabVLA: Targeted Backdoor Attacks on Vision-Language-Action Models
- More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
- UniCoD: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
- Align2Act: Instruction-Tuned Models for Human-Aligned Autonomous Driving
- ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- Dejavu: Towards Experience Feedback Learning for Embodied Intelligence
- Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
- Fundamentals of Building Autonomous LLM Agents
- Goal-oriented Backdoor Attack against Vision-Language-Action Models via Physical Objects
- R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- USIM and U0: A Vision-Language-Action Dataset and Model for General Underwater Robots
- Executable Analytic Concepts as the Missing Link Between VLM Insight and Precise Manipulation
- BLAZER: Bootstrapping LLM-based Manipulation Agents with Zero-Shot Data Generation
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- FLEET: Formal Language-Grounded Scheduling for Heterogeneous Robot Teams
- ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL Problems
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Bring the Apple, Not the Sofa: Impact of Irrelevant Context in Embodied AI Commands on VLA Models
- FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
- VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
- Medical Vision Language Models as Policies for Robotic Surgery
- INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models
- VENTURA: Adapting Image Diffusion Models for Unified Task Conditioned Navigation
- MetaVLA: Unified Meta Co-training For Efficient Embodied Adaption
- FORGE-Tree: Diffusion-Forcing Tree Search for Long-Horizon Robot Manipulation
- ARRC: Advanced Reasoning Robot Control - Knowledge-Driven Autonomous Manipulation Using Retrieval-Augmented Generation
- EmbodiedCoder: Parameterized Embodied Mobile Manipulation via Modern Coding Model
- D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
- Verifier-free Test-Time Sampling for Vision-Language-Action Models
- Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
- Exploring OCR-augmented Generation for Bilingual VQA
- LangGrasp: Leveraging Fine-Tuned LLMs for Language Interactive Robot Grasping with Ambiguous Instructions
- Contrastive Representation Regularization for Vision-Language-Action Models
- VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
- ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
- NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
- GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
- HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
- Hybrid Training for Vision-Language-Action Models
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
- LLM-MCoX: Large Language Model-based Multi-robot Coordinated Exploration and Search
- ExoPredicator: Learning Abstract Models of Dynamic Worlds for Robot Planning
- Seeing Space and Motion: Enhancing Latent Actions with Spatial and Dynamic Awareness for VLA
- Reinforced Embodied Planning with Verifiable Reward for Real-World Robotic Manipulation
- Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
- Best of Sim and Real: Decoupled Visuomotor Manipulation via Learning Control in Simulation and Perception in Real
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation
- dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
- Effective Model Pruning
- Data-Efficient Multitask DAgger
- SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation
- AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
- AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation
- Agentic Services Computing
- Fidelity-Aware Data Composition for Robust Robot Generalization
- PhysiAgent: An Embodied Agent Framework in Physical World
- IA-VLA: Input Augmentation for Vision-Language-Action models in settings with semantically complex tasks
- Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
- Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints
- DexFlyWheel: A Scalable and Self-improving Data Generation Framework for Dexterous Manipulation
- Mash, Spread, Slice! Learning to Manipulate Object States via Visual Spatial Progress
- GLUE: Global-Local Unified Encoding for Imitation Learning via Key-Patch Tracking
- LAGEA: Language Guided Embodied Agents for Robotic Manipulation
- Evaluating point-light biological motion in multimodal large language models
- Pixel Motion Diffusion is What We Need for Robot Control
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
- RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
- Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
- Developing Vision-Language-Action Model from Egocentric Videos
- SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
- EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transfer
- Vision Language Models Cannot Plan, but Can They Formalize?
- VC-Agent: An Interactive Agent for Customized Video Dataset Collection
- RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models
- CAD-Tokenizer: Towards Text-based CAD Prototyping via Modality-Specific Tokenization
- LayoutAgent: A Vision-Language Agent Guided Compositional Diffusion for Spatial Layout Planning
- mindmap: Spatial Memory in Deep Feature Maps for 3D Action Policies
- One Filters All: A Generalist Filter for State Estimation
- Embodied AI: From LLMs to World Models
- FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
- Thinking While Listening: Simple Test Time Scaling For Audio Classification
- EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data
- Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
- Score the Steps, Not Just the Goal: VLM-Based Subgoal Evaluation for Robotic Manipulation
- SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
- Cross-Embodiment Transfer via Behavior-Aligned Representations
- LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- TAPO: Transition-Aware Policy Optimization for LLM Agents
- RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy
- Embodied large language models enable robots to complete complex tasks in unpredictable environments
- MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
- SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
- Position: Human-Robot Interaction in Embodied Intelligence Demands a Shift From Static Privacy Controls to Dynamic Learning
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- Eva-VLA: Evaluating Vision-Language-Action Models' Robustness Under Real-World Physical Variations
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- Do You Need Proprioceptive States in Visuomotor Policies?
- VGGT-DP: Generalizable Robot Control via Vision Foundation Models
- Residual Off-Policy RL for Finetuning Behavior Cloning Policies
- SINGER: An Onboard Generalist Vision-Language Navigation Policy for Drones
- Latent Action Pretraining Through World Modeling
- Language-in-the-Loop Culvert Inspection on the Erie Canal
- ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion
- SEBVS: Synthetic Event-based Visual Servoing for Robot Navigation and Manipulation
- Prepare Before You Act: Learning From Humans to Rearrange Initial States
- RoboSeek: You Need to Interact with Your Objects
- nDNA -- the Semantic Helix of Artificial Cognition
- A Reliable Robot Motion Planner in Complex Real-world Environments via Action Imagination
- From Prediction to Understanding: Will AI Foundation Models Transform Brain Science?
- History-Aware Visuomotor Policy Learning via Point Tracking
- KV-Efficient VLA: A Method to Speed up Vision Language Models with RNN-Gated Chunked KV Cache
- TranTac: Leveraging Transient Tactile Signals for Contact-Rich Robotic Manipulation
- CoReVLA: A Dual-Stage End-to-End Autonomous Driving Framework for Long-Tail Scenarios via Collect-and-Refine
- GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation
- Emulating Human-like Adaptive Vision for Efficient and Flexible Machine Visual Perception
- How Good are Foundation Models in Step-by-Step Embodied Reasoning?
- Self-Improving Embodied Foundation Models
- Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue
- ExT: Towards Scalable Autonomous Excavation via Large-Scale Multi-Task Pretraining and Fine-Tuning
- CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human
- Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI
- VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation
- Toward Embodiment Equivariant Vision-Language-Action Policy
- CLAW: A Vision-Language-Action Framework for Weight-Aware Robotic Grasping
- DREAM: Domain-aware Reasoning for Efficient Autonomous Underwater Monitoring
- SeqVLA: Sequential Task Execution for Long-Horizon Manipulation with Completion-Aware Vision-Language-Action Model
- GestOS: Advanced Hand Gesture Interpretation via Large Language Models to control Any Type of Robot
- GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
- PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models
- When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective
- HARMONIC: A Content-Centric Cognitive Robotic Architecture
- Bridging Perception and Planning: Towards End-to-End Planning for Signal Temporal Logic Tasks
- Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
- Igniting VLMs toward the Embodied Space
- AssemMate: Graph-Based LLM for Robotic Assembly Assistance
- Cross-Platform Scaling of Vision-Language-Action Models from Edge to Cloud GPUs
- Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
- ActivePose: Active 6D Object Pose Estimation and Tracking for Robotic Manipulation
- FEWT: Improving Humanoid Robot Perception with Frequency-Enhanced Wavelet-based Transformers
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
- HumanoidVerse: A Versatile Humanoid for Vision-Language Guided Multi-Object Rearrangement
- ZapGPT: Free-form Language Prompting for Simulated Cellular Control
- Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models
- NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
- RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction
- Graph-Fused Vision-Language-Action for Policy Reasoning in Multi-Arm Robotic Manipulation
- Text2Touch: Tactile In-Hand Manipulation with LLM-Designed Reward Functions
- CRISP -- Compliant ROS2 Controllers for Learning-Based Manipulation Policies and Teleoperation
- Towards Trustworthy Agentic IoEV: AI Agents for Explainable Cyberthreat Mitigation and State Analytics
- LLaDA-VLA: Vision Language Diffusion Action Models
- Robotic Manipulation Framework Based on Semantic Keypoints for Packing Shoes of Different Sizes, Shapes, and Softness
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
- FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies
- COMMET: A System for Human-Induced Conflicts in Mobile Manipulation of Everyday Tasks
- EMMA: Scaling Mobile Manipulation via Egocentric Human Data
- Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
- Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning
- FPC-VLA: A Vision-Language-Action Framework with a Supervisor for Failure Prediction and Correction
- Long-Horizon Visual Imitation Learning via Plan and Code Reflection
- ANNIE: Be Careful of Your Robots
- Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
- Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception
- Galaxea Open-World Dataset and G0 Dual-System VLA Model
- Transforming Agency. On the mode of existence of Large Language Models
- Dynamics-Compliant Trajectory Diffusion for Super-Nominal Payload Manipulation
- AI Compute Architecture and Evolution Trends
- EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control
- CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
- Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
- Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Multimodal Data Storage and Retrieval for Embodied AI: A Survey
- CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Improving Pre-Trained Vision-Language-Action Policies with Model-Based Search
- Self-Guided Action Diffusion
- OmniD: Generalizable Robot Manipulation Policy via Image-Based BEV Representation
- Human Centric General Physical Intelligence for Agile Manufacturing Automation
- TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
- Inspire or Predict? Exploring New Paradigms in Assisting Classical Planners with Large Language Models
- Utilizing Vision-Language Models as Action Models for Intent Recognition and Assistance
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- A Semantic-Aware Framework for Safe and Intent-Integrative Assistance in Upper-Limb Exoskeletons
- ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
- Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
- OmniVTLA: Vision-Tactile-Language-Action Model with Semantic-Aligned Tactile Sensing
- DeepFleet: Multi-Agent Foundation Models for Mobile Robots
- Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
- GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
- \(X\)-evolve: Solution space evolution powered by large language models
- AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning
- MolmoAct: Action Reasoning Models that can Reason in Space
- GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
- Triple-S: A Collaborative Multi-LLM Framework for Solving Long-Horizon Implicative Tasks in Robotics
- Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
- ADPro: a Test-time Adaptive Diffusion Policy via Manifold-constrained Denoising and Task-aware Initialization for Robotic Manipulation
- PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- The Missing Reward: Active Inference in the Era of Experience
- DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
- Information-Theoretic Graph Fusion with Vision-Language-Action Model for Policy Reasoning and Dual Robotic Control
- Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
- Analyzing the Impact of Multimodal Perception on Sample Complexity and Optimization Landscapes in Imitation Learning
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
- Learning Robust Intervention Representations with Delta Embeddings
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- A tutorial note on collecting simulated data for vision-language-action models
- CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation
- ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
- Point2Act: Efficient 3D Distillation of Multimodal LLMs for Zero-Shot Context-Aware Grasping
- FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
- RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models
- ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
- VLH: Vision-Language-Haptics Foundation Model
- On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
- HannesImitation: Grasping with the Hannes Prosthetic Hand via Imitation Learning
- Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents
- H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
- Distributed AI Agents for Cognitive Underwater Robot Autonomy
- Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
- Improving Generalization Ability of Robotic Imitation Learning by Resolving Causal Confusion in Observations
Related