Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
2025/05/26 by Lisondra, Matthew, Benhabib, Beno, Nejat, Goldie · 1 citation
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2505.20503
Abstract
Rapid advancements in foundation models, including Large Language Models, Vision-Language Models, Multimodal Large Language Models, and Vision-Language-Action Models have opened new avenues for embodied AI in mobile service robotics. By combining foundation models with the principles of embodied AI, where intelligent systems perceive, reason, and act through physical interactions, robots can improve understanding, adapt to, and execute complex tasks in dynamic real-world environments. However, embodied AI in mobile service robots continues to face key challenges, including multimodal sensor fusion, real-time decision-making under uncertainty, task generalization, and effective human-robot interactions (HRI). In this paper, we present the first systematic review of the integration of foundation models in mobile service robotics, identifying key open challenges in embodied AI and examining how foundation models can address them. Namely, we explore the role of such models in enabling real-time sensor fusion, language-conditioned control, and adaptive task execution. Furthermore, we discuss real-world applications in the domestic assistance, healthcare, and service automation sectors, demonstrating the transformative impact of foundation models on service robotics. We also include potential future research directions, emphasizing the need for predictive scaling laws, autonomous long-term adaptation, and cross-embodiment generalization to enable scalable, efficient, and robust deployment of foundation models in human-centric robotic systems.
Citations
- SplatSearch: Instance Image Goal Navigation for Mobile Robots using 3D Gaussian Splatting and Diffusion Models
- A Human-in-the-loop Approach to Robot Action Replanning through LLM Common-Sense Reasoning
- X-Nav: Learning End-to-End Cross-Embodiment Navigation for Mobile Robots
- SAFE: Multitask Failure Detection for Vision-Language-Action Models
- π0.5: a Vision-Language-Action Model with Open-World Generalization
- Cooking Task Planning using LLM and Verified by Graph Network
- Aether: Geometric-Aware Unified World Modeling
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- DreamerV3 for Traffic Signal Control: Hyperparameter Tuning and Performance
- Magma: A Foundation Model for Multimodal AI Agents
- LIMO: Less is More for Reasoning
- VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
- s1: Simple test-time scaling
- Mobile Robot Navigation Using Hand-Drawn Maps: A Vision Language Model Approach
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- MiniMax-01: Scaling Foundation Models with Lightning Attention
- Cosmos World Foundation Model Platform for Physical AI
- Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability
- The One RING: a Robotic Indoor Navigation Generalist
- TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies
- Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
- CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs
- Multimodal Alignment and Fusion: A Survey
- Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation
- DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
- PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks
- π0: A Vision-Language-Action Flow Model for General Robot Control
- GPT-4o System Card
- Cocoon: Robust Multi-Modal Perception with Uncertainty-Aware Sensor Fusion
- Robi Butler: Multimodal Remote Interaction with a Household Robot Assistant
- GSON: A Group-based Social Navigation Framework with Large Multimodal Model
- FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning
- OLiVia-Nav: An Online Lifelong Vision Language Approach for Mobile Robot Social Navigation
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Moshi: a speech-text foundation model for real-time dialogue
- SAM 2: Segment Anything in Images and Videos
- The Llama 3 Herd of Models
- ReplanVLM: Replanning Robotic Tasks with Visual Language Models
- PartGLEE: A Foundation Model for Recognizing and Parsing Any Objects
- Situated Instruction Following
- PaliGemma: A versatile 3B VLM for transfer
- Mobile Edge Intelligence for Large Language Models: A Contemporary Survey
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- Unveiling Encoder-Free Vision-Language Models
- RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
- OpenVLA: An Open-Source Vision-Language-Action Model
- Octo: An Open-Source Generalist Robot Policy
- Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- LLeMpower: Understanding Disparities in the Control and Access of Large Language Models
- Smart Help: Strategic Opponent Modeling for Proactive and Adaptive Robot Assistance in Households
- InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
- VLM-Social-Nav: Socially Aware Robot Navigation through Scoring using Vision-Language Models
- Learning Human Preferences Over Robot Behavior as Soft Planning Constraints
- InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
- Interactive Continual Learning Architecture for Long-Term Personalization of Home Service Robots
- Towards Measuring and Modeling "Culture" in LLMs: A Survey
- VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
- CultureLLM: Incorporating Cultural Differences into Large Language Models
- MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
- AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
- Lumiere: A Space-Time Diffusion Model for Video Generation
- Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security
- Bridging Language and Action: A Survey of Language-Conditioned Robot Manipulation
- Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
- Foundation Models in Robotics: Applications, Challenges, and the Future
- Harmonic Mobile Manipulation
- Robot Learning in the Era of Foundation Models: A Survey
- RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation
- Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots
- Bootstrap Your Own Skills: Learning to Solve New Tasks with Large Language Model Guidance
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
- Improved Baselines with Visual Instruction Tuning
- The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
- ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
- Dataset for Open Vocabulary Entity Grounding (DOVE-G)
- InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
- LLM-Grounder: Open-Vocabulary 3D Visual Grounding with Large Language Model as an Agent
- Physically Grounded Vision-Language Models for Robotic Manipulation
- Code Llama: Open Foundation Models for Code
- A Survey on Model Compression for Large Language Models
- OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Large language models in medicine
- SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning
- ViNT: A Foundation Model for Visual Navigation
- TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models
- PaLI-X: On Scaling up a Multilingual Vision and Language Model
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- DINOv2: Learning Robust Visual Features without Supervision
- Micrograph segmentations for DDEVD
- Segment Anything
- Vision-Language Models for Vision Tasks: A Survey
- Vision-Language Models for Vision Tasks: A Survey
- A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models
- GPT-4 Technical Report
- Chat with the Environment: Interactive Multimodal Perception Using Large Language Models
- Understanding the Uncertainty Loop of Human-Robot Interaction
- PaLM-E: An Embodied Multimodal Language Model
- LLaMA: Open and Efficient Foundation Language Models
- Action Dynamics Task Graphs for Learning Plannable Representations of Procedural Tasks
- OpenScene: 3D Scene Understanding with Open Vocabularies
- Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks
- VIMA: General Robot Manipulation with Multimodal Prompts
- ProgPrompt: Generating Situated Robot Task Plans using Large Language Models
- Open-vocabulary Queryable Scene Representations for Real World Planning
- Code as Policies: Language Model Programs for Embodied Control
- Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action
- Masked Autoencoders As Spatiotemporal Learners
- OpenAlex Snapshot
- DeiT III: Revenge of the ViT
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- PaLM: Scaling Language Modeling with Pathways
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Auto-scaling Vision Transformers without Training
- EdgeFormer: A Parameter-Efficient Transformer for On-Device Seq2seq Generation
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
- Geometry-based Graph Pruning for Lifelong SLAM
- CLIPort: What and Where Pathways for Robotic Manipulation
- Self-supervised Reinforcement Learning with Independently Controllable Subgoals
- On the Opportunities and Risks of Foundation Models
- Perceiver IO: A General Architecture for Structured Inputs & Outputs
- A Survey of Uncertainty in Deep Neural Networks
- Evaluating Large Language Models Trained on Code
- A Survey on Deep Learning Technique for Video Segmentation
- Habitat 2.0: Training Home Assistants to Rearrange their Habitat
- BEiT: BERT Pre-Training of Image Transformers
- Language Understanding for Field and Service Robots in a Priori Unknown Environments
- The PRISMA 2020 statement: An updated guideline for reporting systematic reviews
- Learning Transferable Visual Models From Natural Language Supervision
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Training data-efficient image transformers & distillation through attention
- Same Object, Different Grasps: Data and Semantic Knowledge for Task-Oriented Grasping
- Mastering Atari with Discrete World Models
- Mind Your Manners! A Dataset and A Continual Learning Approach for Assessing Social Appropriateness of Robot Actions
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy
- Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy
- Dream to Control: Learning Behaviors by Latent Imagination
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- LIC-Fusion: LiDAR-Inertial-Camera Odometry
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- Multimodal Uncertainty Reduction for Intention Recognition in Human-Robot Interaction
- Habitat: A Platform for Embodied AI Research
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation
- The Limits and Potentials of Deep Learning for Robotics
- Training Deep Networks with Synthetic Data: Bridging the Reality Gap by\n Domain Randomization
- Adversarial Training for Adverse Conditions: Robust Metric Localisation using Appearance Transfer
- A survey of robotic motion planning in dynamic environments
- Deep Multimodal Learning: A Survey on Recent Advances and Trends
- Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
- On Calibration of Modern Neural Networks
- Visual Semantic Planning using Deep Successor Representations
- 1 year, 1000 km: The Oxford RobotCar dataset
- Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles
- ElasticFusion: Real-time dense SLAM and light source estimation
- The Option-Critic Architecture
- Benchmarking Deep Reinforcement Learning for Continuous Control
- Deep Residual Learning for Image Recognition
- You Only Look Once: Unified, Real-Time Object Detection
- Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning
- End-to-End Training of Deep Visuomotor Policies
- Planning and acting in partially observable stochastic domains
- Strips: A new approach to the application of theorem proving to problem solving
- Gemini Robotics: Bringing AI into the Physical World
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Cited by
Related