Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
2023/10/06 by Andy Zhou, Kai Yan, Zhou, Andy +7 · 1 voice · 139 citations
Computer Science · Psychology · #AI-based Problem Solving and Planning #Artificial intelligence #Computer science #Generality #Machine learning #Psychology #Reinforcement Learning in Robotics #Reinforcement learning #Topic Modeling #Tree (set theory)
paper · pdf · doi:10.48550/arxiv.2310.04406
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/10/06 · openalex created_date 2023/12/07 · openalex updated_date 2026/08/01
Abstract
While language models (LMs) have shown potential across a range of decision-making tasks, their reliance on simple acting processes limits their broad deployment as autonomous agents. In this paper, we introduce Language Agent Tree Search (LATS) -- the first general framework that synergizes the capabilities of LMs in reasoning, acting, and planning. By leveraging the in-context learning ability of LMs, we integrate Monte Carlo Tree Search into LATS to enable LMs as agents, along with LM-powered value functions and self-reflections for proficient exploration and enhanced decision-making. A key feature of our approach is the incorporation of an environment for external feedback, which offers a more deliberate and adaptive problem-solving mechanism that surpasses the constraints of existing techniques. Our experimental evaluation across diverse domains, including programming, interactive question-answering (QA), web navigation, and math, validates the effectiveness and generality of LATS in decision-making while maintaining competitive or improved reasoning performance. Notably, LATS achieves state-of-the-art pass@1 accuracy (92.7%) for programming on HumanEval with GPT-4 and demonstrates gradient-free performance (average score of 75.9) comparable to gradient-based fine-tuning for web navigation on WebShop with GPT-3.5. Code can be found at https://github.com/lapisrocks/LanguageAgentTreeSearch
Cited by
- MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing
- Copy-on-Write Scoring: Application-Specific Agent Evaluations
- RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning
- From Memory to Skills: Evidence-Grounded Co-Evolution Governance for Long-Horizon LLM Agents
- Trajectory-Aware Retrieval Agents for Temporal Decision- Making
- Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
- Mitigating Conversational Inertia in Multi-Turn Agents
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search
- s1: Simple test-time scaling
- SPIRAL: Symbolic LLM Planning via Grounded and Reflective Search
- DeepLook: Deeper Thinking with Lookahead
- Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines
- A Unified Definition of Hallucination: It's The World Model, Stupid!
- Synthesizing Procedural Memory: Challenges and Architectures in Automated Workflow Generation
- SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
- CangLing-KnowFlow: A Unified Knowledge-and-Flow-fused Agent for Comprehensive Remote Sensing Applications
- WebOperator: Action-Aware Tree Search for Autonomous Agents in Web Environment
- ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
- When AI Agents Compete for Jobs: Strategic Capabilities and Economic Dynamics of AI Labour Markets
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems
- Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
- COACH: Collaborative Agents for Contextual Highlighting -- A Multi-Agent Framework for Sports Video Analysis
- ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
- Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
- RPM-MCTS: Knowledge-Retrieval as Process Reward Model with Monte Carlo Tree Search for Code Generation
- VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
- SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
- Boosting In-Silicon Directed Evolution with Fine-Tuned Protein Language Model and Tree Search
- MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation
- Self-Abstraction from Grounded Experience for Plan-Guided Policy Refinement
- ReAcTree: Hierarchical LLM Agent Trees with Control Flow for Long-Horizon Task Planning
- Continual Learning, Not Training: Online Adaptation For Agents
- Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective
- Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective
- Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry
- LLM-as-a-Verifier: A General-Purpose Verification Framework
- Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents
- Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied Reasoning
- PlanU: Large Language Model Reasoning through Planning under Uncertainty
- Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models
- ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
- Cog-Rethinker: Hierarchical Metacognitive Reinforcement Learning for LLM Reasoning
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
- MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding
- Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
- GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
- JEF-Hinter: Leveraging Offline Knowledge for Improving Web Agents Adaptation
- Learning Efficient Guardrails for Compliance
- JoyAgent-JDGenie: Technical Report on the GAIA
- Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution
- Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models
- HEART: Emotionally-driven test-time scaling of Language Models
- Tree Search for LLM Agent Reinforcement Learning
- Training Task Reasoning LLM Agents for Multi-turn Task Planning via Single-turn Reinforcement Learning
- Reflect before Act: Proactive Error Correction in Language Models
- Can Multi-turn Self-refined Single Agent LMs with Retrieval Solve Hard Coding Problems?
- Transforming Agency. On the mode of existence of Large Language Models
- Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
- Self-Organizing Agent Network for LLM-based Workflow Automation
- Effective Red-Teaming of Policy-Adherent Agents
- Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
- Large Language Model-based Data Science Agent: A Survey
- LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra
- DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer
- Agentic AI for autonomous anomaly management in complex systems
- PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
- Does More Inference-Time Compute Really Help Robustness?
- FCRF: Flexible Constructivism Reflection for Long-Horizon Robotic Task Planning with Large Language Models
- DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback
- DPBench: Structural Determinants of Multi-Agent LLM Coordination Under Simultaneous Resource Contention
- Enhancing Test-Time Scaling of Large Language Models with Hierarchical Retrieval-Augmented MCTS
- APPO: Agentic Procedural Policy Optimization
- Contextual Experience Replay for Self-Improvement of Language Agents
- Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments
- Prover Agent: An Agent-Based Framework for Formal Mathematical Proofs
- HiMA-Ecom: Enabling Joint Training of Hierarchical Multi-Agent E-commerce Assistants
- SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition
- Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities
- Toward Autonomous UI Exploration: The UIExplorer Benchmark
- When Can Model-Free Reinforcement Learning be Enough for Thinking?
- OAgents: An Empirical Study of Building Effective Agents
- DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents
- Can Theoretical Physics Research Benefit from Language Agents?
- LLM-First Search: Self-Guided Exploration of the Solution Space
- Build Agent Advocates, Not Platform Agents
- CyclicReflex: Improving Reasoning Models via Cyclical Reflection Token Scheduling
- Reason from Future: Reverse Thought Chain Enhances LLM Reasoning
- The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective
- Enhancing Decision-Making of Large Language Models via Actor-Critic
- TreeRare: Syntax Tree-Guided Retrieval and Reasoning for Knowledge-Intensive Question Answering
- MIRROR: Converging Cognitive Principles as Computational Mechanisms for AI Reasoning
- Dyna-Think: Synergizing Reasoning, Acting, and World Model Simulation in AI Agents
- Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
- Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent Systems
- HyperTree Planning: Enhancing LLM Reasoning via Hierarchical Thinking
- T2Agent A Tool-augmented Multimodal Misinformation Detection Agent with Monte Carlo Tree Search
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- Multi-Agent Collaboration via Evolving Orchestration
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
- syftr: Pareto-Optimal Generative AI
- Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
- Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL
- TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
- Cost-Awareness in Tree-Search LLM Planning: A Systematic Study
- TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
- Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM
- Navigating the Alpha Jungle: An LLM-Powered MCTS Framework for Formulaic Factor Mining
- Long-Term Memory for VLA-based Agents in Open-World Task Execution
- Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
- AgentSpec: Understanding Embodied Agent Scaffolds Through Controlled Composition
- ParEVO: Synthesizing Code for Irregular Data: High-Performance Parallelism through Agentic Evolution
- Benchmark Test-Time Scaling of General LLM Agents
- Scaling Laws for Agent Harnesses via Effective Feedback Compute
- Agentic Test-Time Scaling for WebAgents
- From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
- OpenSkill: Open-World Self-Evolution for LLM Agents
- Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces
- Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
- On the Role of Computation in Reinforcement Learning
- Vulcan: Instance-specialized, Verifiable Systems Heuristics Through LLM-driven Search
- Tree of Thoughts as a Classical Heuristic Search Problem: Formal Foundations and Design Patterns
- WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
- MR. Video: "MapReduce" is the Principle for Long Video Understanding
- TTRL: Test-Time Reinforcement Learning
- PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
- CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation
- Are Retrials All You Need? Enhancing Large Language Model Reasoning Without Verbalized Feedback
- WebRollback: Enhancing Web Agents with Explicit Rollback Mechanisms
- SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation
- PPDL: LLM-Based Flows as Probabilistic Programs
- Offline Learning and Forgetting for Reasoning with Large Language Models
- Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning
- RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
- LSR-MCTS: Alleviating Long Range Dependency in Code Generation
- To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
- A Desideratum for Conversational Agents: Capabilities, Challenges, and Future Directions
Discussions
Related