Human Centric General Physical Intelligence for Agile Manufacturing Automation
2025/08/16 by Sandeep Kanta, Mehrdad Tavassoli, Kanta, Sandeep +10 · 1 citation
Engineering · #Digital Transformation in Industry #FOS: Computer and information sciences #Flexible and Reconfigurable Manufacturing Systems #Robotics (cs.RO)
paper · pdf · doi:10.48550/arxiv.2508.11960
openalex publication_date 2025/08/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Abstract
Agile human-centric manufacturing increasingly requires resilient robotic solutions that are capable of safe and productive interactions within unstructured environments of modern factories. While multi-modal sensor fusion provides comprehensive situational awareness yet robots must also contextualize their reasoning to achieve deep semantic understanding of complex scenes. Foundation model particularly Vision-Language-Action (VLA) models have emerged as promising approach on integrating diverse perceptual modalities and spatio-temporal reasoning abilities to ground physical actions to realize General Physical Intelligence (GPI) across various robotic embodiments. Although GPI has been conceptually discussed in literature but its pivotal role and practical deployment in agile manufacturing remain underexplored. To address this gap, this practical review systematically surveys recent advances in VLA models through the lens of GPI by offering comparative analysis of leading implementations and evaluating their industrial readiness via structured ablation study. The state of the art is organized into six thematic pillars including multisensory representation learning, sim2real transfer, planning and control, uncertainty and safety measures and benchmarking. Finally, the review highlights open challenges and future directions for integrating GPI into industrial ecosystems to align with the vision of Industry 5.0 for intelligent, adaptive and collaborative manufacturing ecosystem.
Citations
- ReGen: Generative Robot Simulation via Inverse Design
- Reasoning in Space via Grounding in the World
- RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models
- A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding
- A Survey: Learning Embodied Intelligence from Physical Simulators and World Models
- STEP Planner: Constructing cross-hierarchical subgoal tree as an embodied long-horizon task planner
- ROSA: Harnessing Robot States for Vision-Language and Action Alignment
- ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation
- OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
- EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation
- Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
- Towards Cognitive Collaborative Robots: Semantic-Level Integration and Explainable Control for Human-Centric Cooperation
- MOSAIC: Skill-Centric Manipulation Planning with Physics Simulation
- π0.5: a Vision-Language-Action Model with Open-World Generalization
- Aligning Diffusion Model with Problem Constraints for Trajectory Optimization
- Learning 3D Object Spatial Relationships from Pre-trained 2D Diffusion Models
- Safety Aware Task Planning via Large Language Models in Robotics
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
- ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
- Generating Robot Constitutions & Benchmarks for Semantic Safety
- CoinRobot: Generalized End-to-end Robotic Learning for Physical Intelligence
- MuBlE: MuJoCo and Blender simulation Environment and Benchmark for Task Planning in Robot Manipulation
- Code-as-Symbolic-Planner: Foundation Model-Based Robot Planning via Symbolic Code Generation
- Physics-Driven Data Generation for Contact-Rich Manipulation via Trajectory Optimization
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- BOSS: Benchmark for Observation Space Shift in Long-Horizon Task
- Exploring Embodied Multimodal Large Models: Development, Datasets, and Future Directions
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
- HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
- Cosmos World Foundation Model Platform for Physical AI
- CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
- Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
- CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
- RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
- Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
- Know Where You're Uncertain When Planning with Multimodal Foundation Models: A Formal Framework
- SPOT: SE(3) Pose Trajectory Diffusion for Object-Centric Manipulation
- π0: A Vision-Language-Action Flow Model for General Robot Control
- Large Language Models for Manufacturing
- RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
- OrionNav: Online Planning for Robot Autonomy with Context-Aware LLM and Open-Vocabulary Semantic Scene Graphs
- COLLAGE: Collaborative Human-Agent Interaction Generation using Hierarchical Latent Diffusion and Language Models
- Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
- TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
- Benchmarking Sim2Real Gap: High-fidelity Digital Twinning of Agile Manufacturing
- Deep Generative Models in Robotics: A Survey on Learning from Multimodal Demonstrations
- Robotic Control via Embodied Chain-of-Thought Reasoning
- Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
- ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities
- OpenVLA: An Open-Source Vision-Language-Action Model
- SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
- A Survey on Vision-Language-Action Models for Embodied AI
- Octo: An Open-Source Generalist Robot Policy
- Logic-Skill Programming: An Optimization-based Approach to Sequential Skill Planning
- Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks
- Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
- Vision-Language Model-based Physical Reasoning for Robot Liquid Perception
- RT-H: Action Hierarchies Using Language
- RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation
- RAP: Retrieval-Augmented Planning with Contextual Memory for Multimodal LLM Agents
- M2CURL: Sample-Efficient Multimodal Reinforcement Learning via Self-Supervised Representation Learning for Robotic Manipulation
- RePLan: Robotic Replanning with Perception and Language Models
- Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
- Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
- Foundation Models in Robotics: Applications, Challenges, and the Future
- PALM: Predicting Actions through Language Models
- Robot Learning in the Era of Foundation Models: A Survey
- JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models
- Vision-Language Foundation Models as Effective Robot Imitators
- RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation
- Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
- Mastering Robot Manipulation with Multimodal Prompts through Pretraining and Multi-task Fine-tuning
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
- Ferret: Refer and Ground Anything Anywhere at Any Granularity
- MUTEX: Learning Unified Policies from Multimodal Task Specifications
- BridgeData V2: A Dataset for Robot Learning at Scale
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Scaling Up and Distilling Down: Language-Guided Robot Skill Acquisition
- A Two-stage Fine-tuning Strategy for Generalizable Manipulation Skill of Embodied AI
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners
- RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation
- Data-Link: High Fidelity Manufacturing Datasets for Model2Real Transfer under Industrial Settings
- PaLI-X: On Scaling up a Multilingual Vision and Language Model
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought
- Large Language Models as Commonsense Knowledge for Large-Scale Task Planning
- Towards Generalist Robots: A Promising Paradigm via Generative Simulation
- Dynamic distributed decision-making for resilient resource reallocation in disrupted manufacturing systems
- DINOv2: Learning Robust Visual Features without Supervision
- Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?
- Task and Motion Planning with Large Language Models for Object Rearrangement
- PaLM-E: An Embodied Multimodal Language Model
- MimicPlay: Long-Horizon Imitation Learning by Watching Human Play
- ChatGPT for Robotics: Design Principles and Model Abilities
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Human-Timescale Adaptation in an Open-Ended Task Space
- AnyGrasp: Robust and Efficient Grasp Perception in Spatial and Temporal Domains
- RT-1: Robotics Transformer for Real-World Control at Scale
- Learning Skills from Demonstrations: A Trend from Motion Primitives to Experience Abstraction
- Visual Language Maps for Robot Navigation
- VIMA: General Robot Manipulation with Multimodal Prompts
- Code as Policies: Language Model Programs for Embodied Control
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- DayDreamer: World Models for Physical Robot Learning
- Vision- and tactile-based continuous multimodal intention and attention recognition for safer physical human-robot interaction
- A Generalist Agent
- UL2: Unifying Language Learning Paradigms
- Flamingo: a Visual Language Model for Few-Shot Learning
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- RFUniverse: A Multiphysics Simulation Platform for Embodied AI
- CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
- CLIPort: What and Where Pathways for Robotic Manipulation
- Finetuned Language Models Are Zero-Shot Learners
- GelSlim3.0: High-Resolution Measurement of Shape, Force and Slip in a Compact Tactile-Sensing Finger
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Leveraging Kernelized Synergies on Shared Subspace for Precision Grasp and Dexterous Manipulation
- RLBench: The Robot Learning Benchmark & Learning Environment
- Making Sense of Vision and Touch: Learning Multimodal Representations\n for Contact-Rich Tasks
- Causal Confusion in Imitation Learning
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
- Generalized Zero- and Few-Shot Learning via Aligned Variational Autoencoders
- Model compression via distillation and quantization
- Domain Randomization for Transferring Deep Neural Networks from\n Simulation to the Real World
- Gemini Robotics: Bringing AI into the Physical World
Cited by
Related