Graph of Thoughts: Solving Elaborate Problems with Large Language Models
2023/08/18 by Maciej Besta, Nils Blach, Ales Kubicek +10 · 3 voices · 201 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques
paper · pdf · doi:10.1609/aaai.v38i16.29720
Abstract
We introduce Graph of Thoughts (GoT): a framework that advances prompting capabilities in large language models (LLMs) beyond those offered by paradigms such as Chain-of-Thought or Tree of Thoughts (ToT). The key idea and primary advantage of GoT is the ability to model the information generated by an LLM as an arbitrary graph, where units of information ("LLM thoughts") are vertices, and edges correspond to dependencies between these vertices. This approach enables combining arbitrary LLM thoughts into synergistic outcomes, distilling the essence of whole networks of thoughts, or enhancing thoughts using feedback loops. We illustrate that GoT offers advantages over state of the art on different tasks, for example increasing the quality of sorting by 62% over ToT, while simultaneously reducing costs by >31%. We ensure that GoT is extensible with new thought transformations and thus can be used to spearhead new prompting schemes. This work brings the LLM reasoning closer to human thinking or brain mechanisms such as recurrence, both of which form complex networks
Cited by
- Enhancing SLMs for Sustainable Code Optimization in Radio-Astronomy
- Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
- AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
- Theoretical Foundations of max@k Reinforcement Learning
- Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation
- Arbor: A Framework for Reliable Navigation of Critical Conversation Flows
- Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making
- Constrained Path Reasoning: Measuring When Committed Stages Earn Their Cost
- TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning
- Knowledge-Centric Agents for Workflow Generation in ComfyUI
- Coupled Hierarchical Search over Topology and Execution for Agentic Workflow Synthesis
- Superintelligent Retrieval Agent: The Next Frontier of Agentic Retrieval
- The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
- NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code
- HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
- Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems
- A Survey on Large Language Model-based Agents for Statistics and Data Science
- SymStep: Symbolic Step Verification for Logical Reasoning
- Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement
- DeepLook: Deeper Thinking with Lookahead
- Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech
- Reflection Pretraining Enables Token-Level Self-Correction in Biological Sequence Models
- Efficient Mixture-of-Agents Serving via Tree-Structured Routing, Adaptive Pruning, and Dependency-Aware Prefill-Decode Overlap
- When Reasoning Meets Its Laws
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
- Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention
- Cooperative Retrieval-Augmented Generation for Question Answering: Mutual Information Exchange and Ranking by Contrasting Layers
- Architectures for Building Agentic AI
- Enhancing Clinical Note Generation with ICD-10, Clinical Ontology Knowledge Graphs, and Chain-of-Thought Prompting Using GPT-4
- ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
- Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
- JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
- Generative Recursive Reasoning
- Generative AI for Self-Adaptive Systems: State of the Art and Research Roadmap
- On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
- The Art of Scaling Test-Time Compute for Large Language Models
- Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks
- Goal-Oriented Multi-Agent Semantic Networking: Unifying Intents, Semantics, and Intelligence
- Toward a Safe Internet of Agents
- Multi-chain Graph Refinement and Selection for Reliable Reasoning in Large Language Models
- ORION: Teaching Language Models to Reason Efficiently in the Language of Thought
- EWE: An Agentic Framework for Extreme Weather Analysis
- Universe of Thoughts: Enabling Creative Reasoning with Large Language Models
- DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs
- Agint: Agentic Graph Compilation for Software Engineering Agents
- Reasoning With a Star: A Heliophysics Dataset and Benchmark for Agentic Scientific Reasoning
- Efficiency Will Not Lead to Sustainable Reasoning AI
- GPS: General Per-Sample Prompter
- Dynamic Template Selection for Output Token Generation Optimization: MLP-Based and Transformer Approaches
- Trust in Vision-Language Models: Insights from a Participatory User Workshop
- Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO
- Think with Self-Decoupling and Self-Verification: Automated RTL Design with Backtrack-ToT
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- EcoAlign: An Economically Rational Framework for Efficient LVLM Alignment
- Lost in Serialization: Invariance and Generalization of LLM Graph Reasoners
- Mastering Olympiad-Level Physics with Artificial Intelligence
- Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning
- Voice-Interactive Surgical Agent for Multimodal Patient Data Control
- Think Consistently, Reason Efficiently: Energy-Based Calibration for Implicit Chain-of-Thought
- GRAPH-GRPO-LEX: Contract Graph Modeling and Reinforcement Learning with Group Relative Policy Optimization
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
- Modular Task Decomposition and Dynamic Collaboration in Multi-Agent Systems Driven by Large Language Models
- LTD-Bench: Evaluating Large Language Models by Letting Them Draw
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- Using Span Queries to Optimize for Cache and Attention Locality
- TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models
- Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph
- Serve Programs, Not Prompts
- LLM-as-a-Verifier: A General-Purpose Verification Framework
- RAPID: An Efficient Reinforcement Learning Algorithm for Small Language Models
- ReCAP: Recursive Context-Aware Reasoning and Planning for Large Language Model Agents
- AutoStreamPipe: LLM Assisted Automatic Generation of Data Stream Processing Pipelines
- Modeling Hierarchical Thinking in Large Reasoning Models
- You Don't Need Prompt Engineering Anymore: The Prompting Inversion
- CGoT: A Novel Inference Mechanism for Embodied Multi-Agent Systems Using Composable Graphs of Thoughts
- Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
- Co-Sight: Enhancing LLM-Based Agents via Conflict-Aware Meta-Verification and Trustworthy Reasoning with Structured Facts
- Magellan: Guided MCTS for Latent Space Exploration and Novelty Generation
- The Shape of Reasoning: Topological Analysis of Reasoning Traces in Large Language Models
- Code-enabled language models can outperform reasoning models on diverse tasks
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied Reasoning
- DelvePO: Direction-Guided Self-Evolving Framework for Flexible Prompt Optimization
- Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to Them
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Certified Self-Consistency: Statistical Guarantees and Test-Time Training for Reliable Reasoning in LLMs
- TrustResearcher: Automating Knowledge-Grounded and Transparent Research Ideation with Multi-Agent Collaboration
- Prompt Optimization via Retrieved Reasoning Assets and Multi-Agent Analysis
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
- Orchestrating Human-AI Teams: The Manager Agent as a Unifying Research Challenge
- Where to Search: Measure the Prior-Structured Search Space of LLM Agents
- Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey
- LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
- Cog-Rethinker: Hierarchical Metacognitive Reinforcement Learning for LLM Reasoning
- Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
- Scaling Long-Horizon LLM Agent via Context-Folding
- Limits of Emergent Reasoning of Large Language Models in Agentic Frameworks for Deterministic Games
- MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
- SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
- Fundamentals of Building Autonomous LLM Agents
- FOR-Prompting: From Objection to Revision via an Asymmetric Prompting Protocol
- MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding
- Fortifying LLM-Based Code Generation with Graph-Based Reasoning on Secure Coding Practices
- AMAS: Adaptively Determining Communication Topology for LLM-based Multi-Agent System
- Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
- RareAgent: Self-Evolving Reasoning for Drug Repurposing in Rare Diseases
- InvThink: Premortem Reasoning for Safer Language Models
- Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
- DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
- Searching Meta Reasoning Skeleton to Guide LLM Reasoning
- A global log for medical AI
- Lateral Tree-of-Thoughts Surpasses ToT by Incorporating Logically-Consistent, Low-Utility Candidates
- Rethinking Thinking Tokens: LLMs as Improvement Operators
- Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
- Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs
- Where LLM Agents Fail and How They can Learn From Failures
- Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution
- LatentEvolve: Self-Evolving Test-Time Scaling in Latent Space
- SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
- Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm
- Fast Thinking for Large Language Models
- Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
- MedCritical: Enhancing Medical Reasoning in Small Language Models via Self-Collaborative Correction
- HEART: Emotionally-driven test-time scaling of Language Models
- UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning
- R-Capsule: Compressing High-Level Plans for Efficient Large Language Model Reasoning
- Reinforcement Learning-Guided Chain-of-Draft for Token-Efficient Code Generation
- A2R: An Asymmetric Two-Stage Reasoning Framework for Parallel Reasoning
- Teaching Transformers to Solve Combinatorial Problems through Efficient Trial & Error
- Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
- Why Chain of Thought Fails in Clinical Text Understanding
- Eigen-1: Adaptive Multi-Agent Refinement with Monitor-Based RAG for Scientific Reasoning
- StyleBench: Evaluating thinking styles in Large Language Models
- TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix
- CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering
- Federation of Agents: A Semantics-Aware Communication Fabric for Large-Scale Agentic AI
- LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning
- What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
- OptGraph: Large Language Models Enhanced Evolutionary Optimization Via Graph Retrieval-Augmented Generation
- From Found to Designed: Concepts as a Design Axis for Large Language Models
- Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- LASAR: Latent Adaptive Semantic Aligned Reasoning for Generative Recommendation
- Impromptu: a framework for model-driven prompt engineering
- Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph Inference
- MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents
- Evaluating the Safety and Skill Reasoning of Large Reasoning Models Under Compute Constraints
- OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System
- From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning
- Large Language Models as End-to-end Combinatorial Optimization Solvers
- RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
- Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds
- Empathy-R1: A Chain-of-Empathy and Reinforcement Learning Framework for Long-Form Mental Health Support
- THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
- VerilogMonkey: Exploring Parallel Scaling for Automated Verilog Code Generation with LLMs
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- The Few-shot Dilemma: Over-prompting Large Language Models
- LLMAP: LLM-Assisted Multi-Objective Route Planning with User Preferences
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- Who Decides How Knowing Becomes Doing? Redistributing Authority in Human-AI Music Co-Creation
- MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AI
- Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
- Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing
- From Long to Short: LLMs Excel at Trimming Own Reasoning Chains
- Cross-Question Method Reuse in Large Language Models: From Word-Level Prediction to Rational Logical-Layer Reasoning
- Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- MMAPG: A Training-Free Framework for Multimodal Multi-hop Question Answering via Adaptive Planning Graphs
- Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
- Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
- From Canonical to Complex: Benchmarking LLM Capabilities in Undergraduate Thermodynamics
- STARec: An Efficient Agent Framework for Recommender Systems via Autonomous Deliberate Reasoning
- Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
- Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
- The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases
- mSCoRe: a Multilingual and Scalable Benchmark for Skill-based Commonsense Reasoning
- A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report Generation
- A Survey of Optimization Modeling Meets LLMs: Progress and Future Directions
- DySK-Attn: A Framework for Efficient, Real-Time Knowledge Updating in Large Language Models via Dynamic Sparse Knowledge Attention
- Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context
- Deliberative Reasoning Network: An Uncertainty-Driven Paradigm for Belief-Tracked Inference with Pretrained Language Models
- KG-Augmented Executable CoT for Mathematical Coding
- A DbC Inspired Neurosymbolic Layer for Trustworthy Agent Design
- LaTCoder: Converting Webpage Design to Code with Layout-as-Thought
- NeuroSync: Intent-Aware Code-Based Problem Solving via Direct LLM Understanding Modification
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
- TripTailor: A Real-World Benchmark for Personalized Travel Planning
- Blueprint First, Model Second: A Framework for Deterministic LLM Workflow
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
- DynaSwarm: Dynamically Graph Structure Selection for LLM-based Multi-agent System
- Graph-Augmented Large Language Model Agents: Current Progress and Future Prospects
- Cognitive Chain-of-Thought: Structured Multimodal Reasoning about Social Situations
- Adaptive Cluster Collaborativeness Boosts LLMs Medical Decision Support Capacity
- Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
- TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
- Decoupling Knowledge and Reasoning in LLMs: An Exploration Using Cognitive Dual-System Theory
- Enhanced Mycelium of Thought (EMoT): A Bio-Inspired Hierarchical Reasoning Architecture with Strategic Dormancy and Mnemonic Encoding
Discussions
Related