Large language model-based task planning for service robots: A review
2025/10/27 by Bian, Shaohan, Zhang, Ying, Tian, Guohui +4
#FOS: Computer and information sciences #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2510.23357
Abstract
With the rapid advancement of large language models (LLMs) and robotics, service robots are increasingly becoming an integral part of daily life, offering a wide range of services in complex environments. To deliver these services intelligently and efficiently, robust and accurate task planning capabilities are essential. This paper presents a comprehensive overview of the integration of LLMs into service robotics, with a particular focus on their role in enhancing robotic task planning. First, the development and foundational techniques of LLMs, including pre-training, fine-tuning, retrieval-augmented generation (RAG), and prompt engineering, are reviewed. We then explore the application of LLMs as the cognitive core-`brain'-of service robots, discussing how LLMs contribute to improved autonomy and decision-making. Furthermore, recent advancements in LLM-driven task planning across various input modalities are analyzed, including text, visual, audio, and multimodal inputs. Finally, we summarize key challenges and limitations in current research and propose future directions to advance the task planning capabilities of service robots in complex, unstructured domestic environments. This review aims to serve as a valuable reference for researchers and practitioners in the fields of artificial intelligence and robotics.
Citations
- LLM-Enhanced Rapid-Reflex Async-Reflect Embodied Agent for Real-Time Decision-Making in Dynamically Changing Environments
- CLTP: Contrastive Language-Tactile Pre-training for 3D Contact Geometry Understanding
- Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
- A Survey of mmWave Backscatter: Applications, Platforms, and Technologies
- TLA: Tactile-Language-Action Model for Contact-Rich Manipulation
- Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies
- 3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning
- Large Language Models for Multi-Robot Systems: A Survey
- ZISVFM: Zero-Shot Object Instance Segmentation in Indoor Robotic Environments with Vision Foundation Models
- VeriGraph: Scene Graphs for Execution Verifiable Robot Planning
- MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
- RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction
- ReplanVLM: Replanning Robotic Tasks with Visual Language Models
- The Llama 3 Herd of Models
- GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
- Multi-Head RAG: Solving Multi-Aspect Problems with LLMs
- SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
- LLM-based Robot Task Planning with Exceptional Handling for General Purpose Service Robots
- Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
- IM-RAG: Multi-Round Retrieval-Augmented Generation Through Learning Inner Monologues
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DELTA: Decomposed Efficient Long-Term Robot Task Planning using Large Language Models
- Towards Understanding Convergence and Generalization of AdamW
- Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Grasp, See, and Place: Efficient Unknown Object Rearrangement with Policy Structure Prior
- LLMBind: A Unified Modality-Task Integration Framework
- Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
- ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
- A Survey on Data Selection for LLM Instruction Tuning
- Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
- Binding Touch to Everything: Learning Unified Multimodal Tactile Representations
- Large Language Models for Robotics: Opportunities, Challenges, and Perspectives
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Interactive Visual Task Learning for Robots
- A Survey of Reasoning with Foundation Models
- Foundation Models in Robotics: Applications, Challenges, and the Future
- Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI
- Large Language Models for Robotics: A Survey
- Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
- Don't Make Your LLM an Evaluation Benchmark Cheater
- Vision-Language Interpreter for Robot Task Planning
- Chainpoll: A high efficacy method for LLM hallucination detection
- CareerX: A Retrieval-Augmented Generation Framework for Personalized AI-Driven Career Guidance
- SALM: Speech-augmented Language Model with In-context Learning for Speech Recognition and Translation
- Understanding the Effects of RLHF on LLM Generalisation and Diversity
- Improved Baselines with Visual Instruction Tuning
- SMART-LLM: Smart Multi-Agent Robot Task Planning using Large Language Models
- GRID: Scene-Graph-based Instruction-driven Robotic Task Planning
- Can Whisper perform speech-based in-context learning?
- Instruction Tuning for Large Language Models: A Survey
- Embodied Task Planning with Large Language Models
- Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- AudioPaLM: A Large Language Model That Can Speak and Listen
- CLARA: Classifying and Disambiguating User Commands for Reliable Interactive Robotic Agents
- Full Parameter Fine-tuning for Large Language Models with Limited Resources
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
- Robot Task Planning Based on Large Language Model Representing Knowledge with Directed Graph Structures
- MIMIC-IT: Multi-Modal In-Context Instruction Tuning
- Pengi: An Audio Language Model for Audio Tasks
- Generalized Planning in PDDL Domains with Pretrained Large Language Models
- Listen, Think, and Understand
- PaLM 2 Technical Report
- Otter: A Multi-Modal Model with In-Context Instruction Tuning
- mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
- LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Tool Learning with Foundation Models
- Visual Instruction Tuning
- LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
- MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
- GPT-4 Technical Report
- Task and Motion Planning with Large Language Models for Object Rearrangement
- Exphormer: Sparse Transformers for Graphs
- Foundation Models for Decision Making: Problems, Methods, and Opportunities
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Re3: Generating Longer Stories With Recursive Reprompting and Revision
- Visual Language Maps for Robot Navigation
- GLM-130B: An Open Bilingual Pre-trained Model
- ProgPrompt: Generating Situated Robot Task Plans using Large Language Models
- AudioLM: a Language Modeling Approach to Audio Generation
- Neural-Guided RuntimePrediction of Planners for Improved Motion and Task Planning with Graph Neural Networks
- LaKo: Knowledge-driven Visual Question Answering via Late Knowledge-to-Text Injection
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- Sequential Manipulation Planning on Scene Graph
- Deep Learning Approaches to Grasp Synthesis: A Review
- Large Language Models are Zero-Shot Reasoners
- InDistill: Information flow-preserving knowledge distillation for model compression
- Generalizable Task Planning through Representation Pretraining
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
- Flamingo: a Visual Language Model for Few-Shot Learning
- Training language models to follow instructions with human feedback
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
- Wav2CLIP: Learning Robust Audio Representations From CLIP
- Reframing Instructional Prompts to GPTk's Language
- AudioCLIP: Extending CLIP to Image, Text and Audio
- LoRA: Low-Rank Adaptation of Large Language Models
- Core Challenges of Social Robot Navigation: A Survey
- Learning Transferable Visual Models From Natural Language Supervision
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Language Models are Few-Shot Learners
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- SpanBERT: Improving Pre-training by Representing and Predicting Spans
- Unified Language Model Pre-training for Natural Language Understanding and Generation
- Generating Long Sequences with Sparse Transformers
- Parameter-Efficient Transfer Learning for NLP
- An Empirical Model of Large-Batch Training
- Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
- Attention Is All You Need
- AI Benchmark Half-Life in Recursive Corpora: A Theory of Validity Decay under Semantic Leakage and Regeneration
- CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Cited by
Related