Understanding the planning of LLM agents: A survey
2024/02/05 by Xu Huang, Huang, Xu, Weiwen Liu +15 · 1 voice · 105 citations
Computer Science · #Multi-Agent Systems and Negotiation #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2402.02716
Abstract
As Large Language Models (LLMs) have shown significant intelligence, the progress to leverage LLMs as planning modules of autonomous agents has attracted more attention. This survey provides the first systematic view of LLM-based agents planning, covering recent works aiming to improve planning ability. We provide a taxonomy of existing works on LLM-Agent planning, which can be categorized into Task Decomposition, Plan Selection, External Module, Reflection and Memory. Comprehensive analyses are conducted for each direction, and further challenges for the field of research are discussed.
Cited by
- SymStep: Symbolic Step Verification for Logical Reasoning
- From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
- NEMO-4-PAYPAL: Leveraging NVIDIA's Nemo Framework for empowering PayPal's Commerce Agent
- One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents
- Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection
- TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
- Your plan may succeed, but what about failure? Investigating how people use ChatGPT for long-term life task planning
- CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
- Enhancing Agentic RL with Progressive Reward Shaping and Value-based Sampling Policy Optimization
- Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement Learning
- SEAL: Self-Evolving Agentic Learning for Conversational Question Answering over Knowledge Graphs
- Network Self-Configuration based on Fine-Tuned Small Language Models
- NetDeTox: Adversarial and Efficient Evasion of Hardware-Security GNNs via RL-LLM Orchestration
- Toward a Safe Internet of Agents
- Conversational No-code, Multi-agentic Disease Module Identification and Drug Repurposing Prediction with ChatDRex
- Subgoal Graph-Augmented Planning for LLM-Guided Open-World Reinforcement Learning
- VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
- MagicWand: A Universal Agent for Generation and Evaluation Aligned with User Preference
- Towards a General Framework for HTN Modeling with LLMs
- The Belief-Desire-Intention Ontology for modelling mental reality and agency
- Structured Uncertainty guided Clarification for LLM Agents
- Procedural Knowledge Improves Agentic LLM Workflows
- AgentSUMO: An Agentic Framework for Interactive Simulation Scenario Generation in SUMO via Large Language Models
- Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories
- Standardization of Psychiatric Diagnoses -- Role of Fine-tuned LLM Consortium and OpenAI-gpt-oss Reasoning LLM Enabled Decision Support System
- Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
- ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem
- Integrating Machine Learning into Belief-Desire-Intention Agents: Current Advances and Open Challenges
- Human-Agent Collaborative Paper-to-Page Crafting
- Model Context Contracts - MCP-Enabled Framework to Integrate LLMs With Blockchain Smart Contracts
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models
- DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
- DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping
- GenCellAgent: Generalizable, Training-Free Cellular Image Segmentation via Large Language Model Agents
- GOAT: A Training Framework for Goal-Oriented Agent with Tools
- TraceAegis: Securing LLM-Based Agents via Hierarchical and Behavioral Anomaly Detection
- PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
- AwareCompiler: Agentic Context-Aware Compiler Optimization via a Synergistic Knowledge-Data Driven Framework
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- Self-Sovereign Agent
- Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
- Preference-Aware Memory Update for Long-Term LLM Agents
- Fundamentals of Building Autonomous LLM Agents
- COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
- Large Language Models Meet Virtual Cell: A Survey
- Banking Done Right: Redefining Retail Banking with Language-Centric AI
- ReInAgent: A Context-Aware GUI Agent Enabling Human-in-the-Loop Mobile Task Navigation
- Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
- A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks
- From Principles to Practice: A Systematic Study of LLM Serving on Multi-core NPUs
- When Should Users Check? A Decision-Theoretic Model of Confirmation Frequency in Multi-Step AI Agent Tasks
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
- MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
- SDA-PLANNER: State-Dependency Aware Adaptive Planner for Embodied Task Planning
- AgenticIQA: An Agentic Framework for Adaptive and Interpretable Image Quality Assessment
- Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs
- AIPOM: Agent-aware Interactive Planning for Multi-Agent Systems
- Collaborative and Proactive Management of Task-Oriented Conversations
- ToMPO: Training LLM Strategic Decision Making from a Multi-Agent Perspective
- Agentic AutoSurvey: Let LLMs Survey LLMs
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
- AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
- ChemOrch: Empowering LLMs with Chemical Intelligence via Synthetic Instructions
- Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- ClearFairy: Capturing Creative Workflows through Decision Structuring, In-Situ Questioning, and Rationale Inference
- MACO: A Multi-Agent LLM-Based Hardware/Software Co-Design Framework for CGRAs
- H2R: Hierarchical Hindsight Reflection for Multi-Task LLM Agents
- AgenticIE: An Adaptive Agent for Information Extraction from Complex Regulatory Documents
- A Survey of Reasoning and Agentic Systems in Time Series with Large Language Models
- LLMAP: LLM-Assisted Multi-Objective Route Planning with User Preferences
- Virtual Agent Economies
- Population-Aligned Persona Generation for LLM-based Social Simulation
- Towards Fully Automated Molecular Simulations: Multi-Agent Framework for Simulation Setup and Force Field Extraction
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
- EnvX: Agentize Everything with Agentic AI
- VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
- TableMind: An Autonomous Programmatic Agent for Tool-Augmented Table Reasoning
- Code2MCP: Transforming Code Repositories into MCP Services
- Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- Web Fraud Attacks Against LLM-Driven Multi-Agent Systems
- ERank: Fusing Supervised Fine-Tuning and Reinforcement Learning for Effective and Efficient Text Reranking
- Transforming Agency. On the mode of existence of Large Language Models
- SHERPA: A Model-Driven Framework for Large Language Model Execution
- EPIC: Generative AI Platform for Accelerating HPC Operational Data Analytics
- Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning
- HiPlan: Hierarchical Planning for LLM-Based Agents with Adaptive Global-Local Guidance
- Explicit v.s. Implicit Memory: Exploring Multi-hop Complex Reasoning Over Personalized Information
- GTool: Graph Enhanced Tool Planning with Large Language Model
- Standardization of Neuromuscular Reflex Analysis -- Role of Fine-Tuned Vision-Language Model Consortium and OpenAI gpt-oss Reasoning LLM Enabled Decision Support System
- Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework
- Towards Reliable Multi-Agent Systems for Marketing Applications via Reflection, Memory, and Planning
- Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory
- Planning Agents on an Ego-Trip: Leveraging Hybrid Ego-Graph Ensembles for Improved Tool Retrieval in Enterprise Task Planning
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?
- AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
- Pro2Guard: Proactive Runtime Enforcement of LLM Agent Safety via Probabilistic Model Checking
- Evaluation and Benchmarking of LLM Agents: A Survey
- Graph-Augmented Large Language Model Agents: Current Progress and Future Prospects
- Efficient Agents: Building Effective Agents While Reducing Cost
Discussions
Related