Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction
2025/08/07 by Salimpour, Sahar, Fu, Lei, Rachwał, Kajetan +8 · 5 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2508.05294
Abstract
Foundation models, including large language models (LLMs) and vision-language models (VLMs), have recently enabled novel approaches to robot autonomy and human-robot interfaces. In parallel, vision-language-action models (VLAs) or large behavior models (LBMs) are increasing the dexterity and capabilities of robotic systems. This survey paper reviews works that advance agentic applications and architectures, including initial efforts with GPT-style interfaces and more complex systems where AI agents function as coordinators, planners, perception actors, or generalist interfaces. Such agentic architectures allow robots to reason over natural language instructions, invoke APIs, plan task sequences, or assist in operations and diagnostics. In addition to peer-reviewed research, due to the fast-evolving nature of the field, we highlight and include community-driven projects, ROS packages, and industrial frameworks that show emerging trends. We propose a taxonomy for classifying model integration approaches and present a comparative analysis of the role that agents play in different solutions in today's literature.
Citations
- ROSBag MCP Server: Analyzing Robot Data with LLMs for Agentic Embodied AI Applications
- From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding
- Contemporary Agent Technology: LLM-Driven Advancements vs Classic Multi-Agent Systems
- Embodied AI: Emerging Risks and Opportunities for Policy Action
- Distributed AI Agents for Cognitive Underwater Robot Autonomy
- General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting
- ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
- AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges
- RAI: Flexible Agent Framework for Embodied AI
- Multi-agent Embodied AI: Advances and Future Directions
- A survey of agent interoperability protocols: Model Context Protocol (MCP), Agent Communication Protocol (ACP), Agent-to-Agent Protocol (A2A), and Agent Network Protocol (ANP)
- LLM-Empowered Embodied Agent for Memory-Augmented Task Planning in Household Robotics
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation
- M-LLM Based Video Frame Selection for Efficient Video Understanding
- DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
- π0: A Vision-Language-Action Flow Model for General Robot Control
- Enabling Novel Mission Operations and Interactions with ROSA: The Robot Operating System Agent
- BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation
- ConceptAgent: LLM-Driven Precondition Grounding and Tree Search for Robust Task Planning and Execution
- LaMMA-P: Generalizable Multi-Agent Long-Horizon Task Allocation and Planning with LM-Driven PDDL Planner
- SELP: Generating Safe and Efficient Task Plans for Robot Agents with Large Language Models
- Deep Generative Models in Robotics: A Survey on Learning from Multimodal Demonstrations
- Odyssey: Empowering Minecraft Agents with Open-World Skills
- Enhancing Robot Explanation Capabilities through Vision-Language Models: a Preliminary Study by Interpreting Visual Inputs for Improved Human-Robot Interaction
- Large Language Models for Orchestrating Bimanual Robots
- AIOS: LLM Agent Operating System
- AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
- Foundation Models in Robotics: Applications, Challenges, and the Future
- Robot Learning in the Era of Foundation Models: A Survey
- An Embodied Generalist Agent in 3D World
- Large Language Models for Robotics: A Survey
- TD-MPC2: Scalable, Robust World Models for Continuous Control
- SMART-LLM: Smart Multi-Agent Robot Task Planning using Large Language Models
- RoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- ChatGPT for Robotics: Design Principles and Model Abilities
- Mastering Diverse Domains through World Models
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- Code as Policies: Language Model Programs for Embodied Control
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Reinforcement Learning with Augmented Data
- World Models
- Time-Contrastive Networks: Self-Supervised Learning from Video
- End to End Learning for Self-Driving Cars
- Gemini Robotics: Bringing AI into the Physical World
Cited by
Related