Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
2023/04/23 by Tony Z. Zhao, Vikas Kumar, Vikash Kumar +6 · 5 voices · 360 citations
Computer Science · Engineering · #Human Pose and Action Recognition #Reinforcement Learning in Robotics #Robot Manipulation and Learning #cs.LG #cs.RO
paper · pdf · doi:10.48550/arxiv.2304.13705
Abstract
Fine manipulation tasks, such as threading cable ties or slotting a battery, are notoriously difficult for robots because they require precision, careful coordination of contact forces, and closed-loop visual feedback. Performing these tasks typically requires high-end robots, accurate sensors, or careful calibration, which can be expensive and difficult to set up. Can learning enable low-cost and imprecise hardware to perform these fine manipulation tasks? We present a low-cost system that performs end-to-end imitation learning directly from real demonstrations, collected with a custom teleoperation interface. Imitation learning, however, presents its own challenges, particularly in high-precision domains: errors in the policy can compound over time, and human demonstrations can be non-stationary. To address these challenges, we develop a simple yet novel algorithm, Action Chunking with Transformers (ACT), which learns a generative model over action sequences. ACT allows the robot to learn 6 difficult tasks in the real world, such as opening a translucent condiment cup and slotting a battery with 80-90% success, with only 10 minutes worth of demonstrations. Project website: https://tonyzhaozh.github.io/aloha/
Cited by
- ArmnetBench v0.1: Parallel Real-World Evaluation of Manipulation Policies on a Low-Cost Arm Farm
- τ: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- FutureRTC: Real-Time Robot Execution with Anticipatory-Conditioned Action Chunking
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- N0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
- Try Once, Then Optimal: De-Redundified Procedure Memory for Cross-Episode Exploration Amortization
- KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation
- Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
- Decompose and Reorganize: Planning with Primitives and Visuomotor Policies Learned from Demonstrations
- Teaching Tiny VLA Models Where to Look and How to Move
- MEMENTO: Memory-Guided Memetic Code-as-Policy Evolution
- Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features
- Compositional Motion Generation from Demonstration with Object-Centric Neural Fields
- A Visuo-Tactile Data Collection System with Haptic Feedback for Coarse-to-Fine Imitation Learning
- Imitation Learning for Robot Assistance in Open Surgery: A Multi-Policy Evaluation on Suture Following
- RoboCade: Gamifying Robot Data Collection
- UniTacHand: Unified Spatio-Tactile Representation for Human to Robotic Hand Skill Transfer
- Learning Skills from Action-Free Videos
- MaP-AVR: A Meta-Action Planner for Agents Leveraging Vision Language Models and Retrieval-Augmented Generation
- OMP: One-step Meanflow Policy with Directional Alignment
- Point What You Mean: Visually Grounded Instruction Policy
- Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
- SurgiPose: Estimating Surgical Tool Kinematics from Monocular Video for Surgical Robot Learning
- Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- Unifying Deep Predicate Invention with Pre-trained Foundation Models
- AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning
- Kinematics-Aware Diffusion Policy with Consistent 3D Observation and Action Space for Whole-Arm Robotic Manipulation
- ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
- CaFe-TeleVision: A Coarse-to-Fine Teleoperation System with Immersive Situated Visualization for Enhanced Ergonomics
- DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
- AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis
- Universal Dexterous Functional Grasping via Demonstration-Editing Reinforcement Learning
- Post-Training and Test-Time Scaling of Generative Agent Behavior Models for Interactive Autonomous Driving
- Motus: A Unified Latent Action World Model
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- Towards Accessible Physical AI: LoRA-Based Fine-Tuning of VLA Models for Real-World Robot Control
- Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation
- Masked Generative Policy for Robotic Control
- OSMO: Open-Source Tactile Glove for Human-to-Robot Skill Transfer
- Delay-Aware Diffusion Policy: Bridging the Observation-Execution Gap in Dynamic Tasks
- ESPADA: Execution Speedup via Semantics Aware Demonstration Data Downsampling for Imitation Learning
- Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge
- Training-Time Action Conditioning for Efficient Real-Time Chunking
- HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
- FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
- Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation
- Coordinated Humanoid Manipulation with Choice Policies
- FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination
- Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
- VAT: Vision Action Transformer by Unlocking Full Representation of ViT
- PerFACT: Motion Policy with LLM-Powered Dataset Synthesis and Fusion Action-Chunking Transformers
- Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
- SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction
- EfficientFlow: Efficient Equivariant Flow Policy Learning for Embodied AI
- Learning Dexterous Manipulation Skills from Imperfect Simulations
- MV-TAP: Tracking Any Point in Multi-View Videos
- Real-World Robot Control by Deep Active Inference With a Temporally Hierarchical World Model
- Much Ado About Noising: Dispelling the Myths of Generative Robotic Control
- GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
- Constant-Time Motion Planning with Manipulation Behaviors
- IGen: Scalable Data Generation for Robot Learning from Open-World Images
- M3A Policy: Mutable Material Manipulation Augmentation Policy through Photometric Re-rendering
- VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
- CycleManip: Enabling Cyclic Task Manipulation via Effective Historical Perception and Understanding
- Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
- MILE: A Mechanically Isomorphic Exoskeleton Data Collection System with Fingertip Visuotactile Sensing for Dexterous Manipulation
- DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
- Dataset Poisoning Attacks on Behavioral Cloning Policies
- Dynamic Test-Time Compute Scaling in Control Policy: Difficulty-Aware Stochastic Interpolant Policy
- ACE-F: A Cross Embodiment Foldable System with Force Feedback for Dexterous Teleoperation
- HumanoidExo: Scalable Whole-Body Humanoid Manipulation via Wearable Exoskeleton
- ShapeForce: Low-Cost Soft Robotic Wrist for Contact-Rich Manipulation
- MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action Generalization
- SpeedAug: Policy Acceleration via Tempo-Enriched Policy and RL Fine-Tuning
- Mixture of Horizons in Action Chunking
- MicCheck: Repurposing Off-the-Shelf Pin Microphones for Easy and Low-Cost Contact Sensing
- Observer-Actor: Active Vision Imitation Learning with Sparse-View Gaussian Splatting
- EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
- L1 Sample Flow for Efficient Visuomotor Learning
- Contact-Rich Robotic Assembly in Construction via Diffusion Policy Learning
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation
- VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
- Unify Robot Actions in Camera Frame
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- Bi-AQUA: Bilateral Control-Based Imitation Learning for Underwater Robot Arms via Lighting-Aware Action Chunking with Transformers
- In-N-On: Scaling Egocentric Manipulation with in-the-wild and on-task Data
- IPR-1: Interactive Physical Reasoner
- An Alignment-Based Approach to Learning Motions from Demonstrations
- HMC: Learning Heterogeneous Meta-Control for Contact-Rich Loco-Manipulation
- Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
- RoboTidy : A 3D Gaussian Splatting Household Tidying Benchmark for Embodied Navigation and Action
- From Power to Precision: Learning Fine-grained Dexterity for Multi-fingered Robotic Hands
- Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
- VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
- Learning Adaptive Neural Teleoperation for Humanoid Robots: From Inverse Kinematics to End-to-End Control
- Learning a Thousand Tasks in a Day
- SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
- ViPRA: Video Prediction for Robot Actions
- From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies
- ViTaMIn-B: A Reliable and Efficient Visuo-Tactile Bimanual Manipulation Interface
- Gentle Manipulation Policy Learning via Demonstrations from VLM Planned Atomic Skills
- VLM-driven Skill Selection for Robotic Assembly Tasks
- EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
- TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
- Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation
- Decomposed Object Manipulation via Dual-Actor Policy
- MoE-DP: An MoE-Enhanced Diffusion Policy for Robust Long-Horizon Robotic Manipulation with Skill Decomposition and Failure Recovery
- X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations
- Real-to-Sim Robot Policy Evaluation with Gaussian Splatting Simulation of Soft-Body Interactions
- Temporal Action Selection for Action Chunking
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- Investigating Robot Control Policy Learning for Autonomous X-ray-guided Spine Procedures
- Learning-based Cooperative Robotic Paper Wrapping: A Unified Control Policy with Residual Force Control
- WorldPlanner: Monte Carlo Tree Search and MPC with Action-Conditioned Visual World Models
- TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
- LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation
- Dexterous Robotic Piano Playing at Scale
- GauDP: Reinventing Multi-Agent Collaboration through Gaussian-Image Synergy in Diffusion Policies
- Improving Robustness to Out-of-Distribution States in Imitation Learning via Deep Koopman-Boosted Diffusion Policy
- EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations
- Learning Generalizable Visuomotor Policy through Dynamics-Alignment
- End-to-End Dexterous Arm-Hand VLA Policies via Shared Autonomy: VR Teleoperation Augmented by Autonomous Hand VLA Policy for Efficient Data Collection
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
- Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- Tri-Manual Visuomotor Imitation Learning of Robot Policies
- MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization
- Embracing Evolution: A Call for Body-Control Co-Design in Embodied Humanoid Robot
- CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
- Flow with the Force Field: Learning 3D Compliant Flow Matching Policies from Force and Demonstration-Guided Simulation Data
- From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence
- PACE: Phase-Aware Chunk Execution for Robot Policies with Action Chunking
- InDex: Empowering VLA Models with Intent-Conditioned Arm-Hand Coordination for Dexterous Manipulation
- Nautilus: From One Prompt to Plug-and-Play Robot Learning
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- A Humanoid Visual-Tactile-Action Dataset for Contact-Rich Manipulation
- Blindfolded Experts Generalize Better: Insights from Robotic Manipulation and Videogames
- Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
- Awakening Facial Emotional Expressions in Human-Robot
- ManiDP: Manipulability-Aware Diffusion Policy for Posture-Dependent Bimanual Manipulation
- ACG: Action Coherence Guidance for Flow-based VLA models
- Sample By Step, Optimize By Chunk: Chunk-Level GRPO For Text-to-Image Generation
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- FieldGen: From Teleoperated Pre-Manipulation Trajectories to Field-Guided Data Generation
- PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning
- MemER: Scaling Up Memory for Robot Control via Experience Retrieval
- MR-UBi: Mixed Reality-Based Underwater Robot Arm Teleoperation System with Reaction Torque Indicator via Bilateral Control
- Dino-Diffusion Modular Designs Bridge the Cross-Domain Gap in Autonomous Parking
- GSWorld: Closed-Loop Photo-Realistic Simulation Suite for Robotic Manipulation
- Learning Affordances at Inference-Time for Vision-Language-Action Models
- Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
- VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting
- Foveated Compression for Immersive Telepresence Visualization
- Iterative Refinement of Flow Policies in Probability Space for Online Reinforcement Learning
- Botany-Bot: Digital Twin Monitoring of Occluded and Underleaf Plant Structures with Gaussian Splats
- RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- RAPID Hand Prototype: Design of an Affordable, Fully-Actuated Biomimetic Hand for Dexterous Teleoperation
- VO-DP: Semantic-Geometric Adaptive Diffusion Policy for Vision-Only Robotic Manipulation
- QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
- RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning
- Open TeleDex: A Hardware-Agnostic Teleoperation System for Imitation Learning based Dexterous Manipulation
- Spatially anchored Tactile Awareness for Robust Dexterous Manipulation
- Generative Models From and For Sampling-Based MPC: A Bootstrapped Approach For Adaptive Contact-Rich Manipulation
- Prescribed Performance Control of Deformable Object Manipulation in Spatial Latent Space
- ALOHA2 Robot Kitchen Application Scenario Reproduction Report
- VLA-0: Building State-of-the-Art VLAs with Zero Modification
- Learning to Grasp Anything by Playing with Random Toys
- Fast Visuomotor Policy for Robotic Manipulation
- Improving Generative Behavior Cloning via Self-Guidance and Adaptive Chunking
- ARMADA: Autonomous Online Failure Detection and Human Shared Control Empower Scalable Real-world Deployment and Adaptation
- SCOOP'D: Learning Mixed-Liquid-Solid Scooping via Sim2Real Generative Policy
- HiMaCon: Discovering Hierarchical Manipulation Concepts from Unlabeled Multi-Modal Data
- DemoHLM: From One Demonstration to Generalizable Humanoid Loco-Manipulation
- RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action Model
- High-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning with Gaussian Splatting
- Fine-Tuning Flow Matching via Maximum Likelihood Estimation of Reconstructions
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- UF-RNN: Real-Time Adaptive Motion Generation Using Uncertainty-Driven Foresight Prediction
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation
- Failure Prediction at Runtime for Generative Robot Policies
- Glovity: Learning Dexterous Contact-Rich Manipulation via Spatial Wrench Feedback Teleoperation System
- Toward Autonomous Soft Robotic Endovascular Navigation via Imitation Learning
- Symskill: Symbol and Skill Co-Invention for Data-Efficient and Real-Time Long-Horizon Manipulation
- Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation
- R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation
- DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
- ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL Problems
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- RLinf-VLA: A Unified and Efficient Framework for VLA+RL Training
- Avi: Action from Volumetric Inference
- Towards Autonomous Tape Handling for Robotic Wound Redressing
- FORGE-Tree: Diffusion-Forcing Tree Search for Long-Horizon Robot Manipulation
- ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
- Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning
- HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
- MobRT: A Digital Twin-Based Framework for Scalable Learning in Mobile Manipulation
- Feedback Matters: Augmenting Autonomous Dissection with Visual and Topological Feedback
- ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
- Seeing the Bigger Picture: 3D Latent Mapping for Mobile Manipulation Policy Learning
- NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
- GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
- Prometheus: Universal, Open-Source Mocap-Based Teleoperation System with Force Feedback for Dataset Collection in Robot Learning
- RTFF: Random-to-Target Fabric Flattening Policy using Dual-Arm Manipulator
- CroSTAta: Cross-State Transition Attention Transformer for Robotic Manipulation
- HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
- SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation
- Annotation-Free One-Shot Imitation Learning for Multi-Step Manipulation Tasks
- MSG: Multi-Stream Generative Policies for Sample-Efficient Robotic Manipulation
- From Code to Action: Hierarchical Learning of Diffusion-VLM Policies
- U-DiT Policy: U-shaped Diffusion Transformers for Robotic Manipulation
- Leave No Observation Behind: Real-time Correction for VLA Action Chunks
- GLUE: Global-Local Unified Encoding for Imitation Learning via Key-Patch Tracking
- FTACT: Force Torque aware Action Chunking Transformer for Pick-and-Reorient Bottle Task
- Liaohe-CobotMagic-PnP: an Imitation Learning Dataset of Intelligent Robot for Industrial Applications
- WoW: Towards a World omniscient World model Through Embodied Interaction
- EgoDemoGen: Novel Egocentric Demonstration Generation Enables Viewpoint-Robust Manipulation
- DemoGrasp: Universal Dexterous Grasping from a Single Demonstration
- Developing Vision-Language-Action Model from Egocentric Videos
- SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
- VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
- ARMimic: Learning Robotic Manipulation from Passive Human Demonstrations in Augmented Reality
- ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation
- Normalizing Flows are Capable Models for Bi-manual Visuomotor Policy
- Large Pre-Trained Models for Bimanual Manipulation in 3D
- mindmap: Spatial Memory in Deep Feature Maps for 3D Action Policies
- Parse-Augment-Distill: Learning Generalizable Bimanual Visuomotor Policies from Single Human Video
- LLM Trainer: Automated Robotic Data Generating via Demonstration Augmentation using LLMs
- FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
- EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- DexDirect: Direct Kinesthetic Arm Guidance for Efficient Dexterous Demonstration Collection
- It's Not Just More Demos: Counterfactual Action Sensitivity Coverage for Data-Efficient Robust Robot Imitation
- Failure Detection for Surgical Robot Imitation Policies via Flow-Matching World Modeling
- Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation
- RMBench: Memory-Dependent Robotic Manipulation Benchmark with Insights into Policy Design
- SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation
- REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement Learning
- MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
- Self-evolved Imitation Learning in Simulated World
- ROPA: Synthetic Robot Pose Generation for RGB-D Bimanual Data Augmentation
- SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
- ManipForce: Force-Guided Policy Learning with Frequency-Aware Representation for Contact-Rich Manipulation
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation
- MV-UMI: A Scalable Multi-View Interface for Cross-Embodiment Learning
- N2M: Bridging Navigation and Manipulation by Learning Pose Preference from Rollout
- Do You Need Proprioceptive States in Visuomotor Policies?
- Generalizable Domain Adaptation for Sim-and-Real Policy Co-Training
- VGGT-DP: Generalizable Robot Control via Vision Foundation Models
- Residual Off-Policy RL for Finetuning Behavior Cloning Policies
- Learning Geometry-Aware Nonprehensile Pushing and Pulling with Dexterous Hands
- Latent Action Pretraining Through World Modeling
- PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
- Learning Dexterous Manipulation with Quantized Hand State
- Prepare Before You Act: Learning From Humans to Rearrange Initial States
- RoboManipBaselines: A Unified Framework for Imitation Learning in Robotic Manipulation across Real and Simulated Environments
- FILIC: Dual-Loop Force-Guided Imitation Learning with Impedance Torque Control for Contact-Rich Manipulation Tasks
- History-Aware Visuomotor Policy Learning via Point Tracking
- No Need for Real 3D: Fusing 2D Vision with Pseudo 3D Representations for Robotic Manipulation Learning
- TranTac: Leveraging Transient Tactile Signals for Contact-Rich Robotic Manipulation
- Right-Side-Out: Learning Zero-Shot Sim-to-Real Garment Reversal
- LodeStar: Long-horizon Dexterity via Synthetic Data Augmentation from Human Demonstrations
- End-to-end RL Improves Dexterous Grasping Policies
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Self-Improving Embodied Foundation Models
- M4Diffuser: Multi-View Diffusion Policy with Manipulability-Aware Control for Robust Mobile Manipulation
- Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI
- Toward Embodiment Equivariant Vision-Language-Action Policy
- Learning to Pick: A Visuomotor Policy for Clustered Strawberry Picking
- VRScout: Towards Real-Time, Autonomous Testing of Virtual Reality Games
- SeqVLA: Sequential Task Execution for Long-Horizon Manipulation with Completion-Aware Vision-Language-Action Model
- Hand Tracking Streamer
- Dense-Jump Flow Matching with Non-Uniform Time Scheduling for Robotic Policies: Mitigating Multi-Step Inference Degradation
- StageACT: Stage-Conditioned Imitation for Robust Humanoid Door Opening
- Gen2Real: Towards Demo-Free Dexterous Manipulation by Harnessing Generated Video
- Safety filtering of robotic manipulation under environment uncertainty: a computational approach
- Robust Online Residual Refinement via Koopman-Guided Dynamics Modeling
- Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
- Beyond Anthropomorphism: Enhancing Grasping and Eliminating a Degree of Freedom by Fusing the Abduction of Digits Four and Five
- 4D Visual Pre-training for Robot Learning
- Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
- Igniting VLMs toward the Embodied Space
- Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
- FEWT: Improving Humanoid Robot Perception with Frequency-Enhanced Wavelet-based Transformers
- Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference-Scoped Exploration
- Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
- NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows
- Input-gated Bilateral Teleoperation: An Easy-to-implement Force Feedback Teleoperation Method for Low-cost Hardware
- Grasp Like Humans: Learning Generalizable Multi-Fingered Grasping from Human Proprioceptive Sensorimotor Integration
- RoboMatch: A Unified Mobile-Manipulation Teleoperation Platform with Auto-Matching Network Architecture for Long-Horizon Tasks
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
- RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction
- CRISP -- Compliant ROS2 Controllers for Learning-Based Manipulation Policies and Teleoperation
- F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
- Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
- FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies
- Imitation Learning Based on Disentangled Representation Learning of Behavioral Characteristics
- Action Chunking with Transformers for Image-Based Spacecraft Guidance and Control
- EMMA: Scaling Mobile Manipulation via Egocentric Human Data
- FPC-VLA: A Vision-Language-Action Framework with a Supervisor for Failure Prediction and Correction
- DEXOP: A Device for Robotic Transfer of Dexterous Human Manipulation
- ANNIE: Be Careful of Your Robots
- The Role of Embodiment in Intuitive Whole-Body Teleoperation for Mobile Manipulation
- Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
- U-ARM : Ultra low-cost general teleoperation interface for robot manipulation
- OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- Exploiting Policy Idling for Dexterous Manipulation
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- AutoRing: Imitation Learning--based Autonomous Intraocular Foreign Body Removal Manipulation with Eye Surgical Robot
- SafeBimanual: Diffusion-based Trajectory Optimization for Safe Bimanual Manipulation
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Taming VR Teleoperation and Learning from Demonstration for Multi-Task Bimanual Table Service Manipulation
- MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence
- CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- Sim-to-Real Dynamic Object Manipulation on Conveyor Systems via Optimization Path Shaping
- Improving Pre-Trained Vision-Language-Action Policies with Model-Based Search
- Self-Guided Action Diffusion
- Multi-Group Equivariant Augmentation for Reinforcement Learning in Robot Manipulation
- 3D FlowMatch Actor: Unified 3D Policy for Single- and Dual-Arm Manipulation
- MLM: Learning Multi-task Loco-Manipulation Whole-Body Control for Quadruped Robot with Arm
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Leveraging OS-Level Primitives for Robotic Action Management
- Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
- BEAVR: Bimanual, multi-Embodiment, Accessible, Virtual Reality Teleoperation System for Robots
- Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT
- GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
- AgentWorld: An Interactive Simulation Platform for Scene Construction and Mobile Robotic Manipulation
- MolmoAct: Action Reasoning Models that can Reason in Space
- GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
- D3P: Dynamic Denoising Diffusion Policy via Reinforcement Learning
- Affordance-Guided Dual-Armed Disassembly Teleoperation for Mating Parts
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
- Constraint-Preserving Data Generation for Visuomotor Policy Learning
- Safety-Aware Imitation Learning via MPC-Guided Disturbance Injection
- CLASS: Contrastive Learning via Action Sequence Supervision for Robot Manipulation
- Learning to Perform Low-Contact Autonomous Nasotracheal Intubation by Recurrent Action-Confidence Chunking with Transformer
- RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
- On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
- TOP: Time Optimization Policy for Stable and Accurate Standing Manipulation with Humanoid Robots
- Video Generators are Robot Policies
- CHILD (Controller for Humanoid Imitation and Live Demonstration): a Whole-Body Humanoid Teleoperation System
- XRoboToolkit: A Cross-Platform Framework for Robot Teleoperation
- H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
- Improving Generalization Ability of Robotic Imitation Learning by Resolving Causal Confusion in Observations
- DISCOVERSE: Efficient Robot Simulation in Complex High-Fidelity Environments
Discussions
- 🤖 New from Stanford, UC Berkeley & Meta: robots mastering fine bimanual tasks—like threading zip ties & slotting batteries—with low-cost hardware and just 10 minutes of demos.  Introducing ACT: Acti [bsky, 3 points, 0 comments]
- ALOHA - a low-cost bimanual robot system that learns fine manipulation tasks—like opening a condiment cup—with 80–90% success using just 10 minutes of human demos.  Paper: arxiv.org/abs/2304.13705 Pr [bsky, 2 points, 0 comments]
- The reason why these tasks are kinda hard to do is that it involves a surprisingly lot of planning in the motion. Supposed you are trying to place an object from point A to B. The planning is simple. [stackexchange, 1 points]
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware [hn, 1 points, 0 comments]
- Precise manipulation was supposed to need expensive hardware. ALOHA disproves that. ACT hits 80-90% on battery slotting from 10min of demos on a <$20k bimanual setup. In 2026, with VLAs everywhere, th [bsky, 0 points, 1 comments]
Related