Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
2022/11/22 by Wenhu Chen, Chen, Wenhu, Xueguang Ma +5 · 11 voices · 308 citations
Computer Science · #Artificial intelligence #Code (set theory) #Computation #Computer science #Consistency (knowledge bases) #Explainable Artificial Intelligence (XAI) #Language model #Online Learning and Analytics #Process (computing) #Programming language #Source code #Theoretical computer science #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2211.12588
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2022/11/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/02
Abstract
Recently, there has been significant progress in teaching language models to perform step-by-step reasoning to solve complex numerical reasoning tasks. Chain-of-thoughts prompting (CoT) is by far the state-of-art method for these tasks. CoT uses language models to perform both reasoning and computation in the multi-step `thought' process. To disentangle computation from reasoning, we propose `Program of Thoughts' (PoT), which uses language models (mainly Codex) to express the reasoning process as a program. The computation is relegated to an external computer, which executes the generated programs to derive the answer. We evaluate PoT on five math word problem datasets (GSM, AQuA, SVAMP, TabMWP, MultiArith) and three financial-QA datasets (FinQA, ConvFinQA, TATQA) for both few-shot and zero-shot setups. Under both few-shot and zero-shot settings, PoT can show an average performance gain over CoT by around 12% across all the evaluated datasets. By combining PoT with self-consistency decoding, we can achieve SoTA performance on all math problem datasets and near-SoTA performance on financial datasets. All of our data and code are released in Github https://github.com/wenhuchen/Program-of-Thoughts
Cited by
- Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models
- NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
- TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
- Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
- PAGE-RAG: Evidence-Grounded Adaptive Graph Retrieval for Long-Document Question Answering
- Gold-Guided Programmatic Distillation for Financial Reasoning over Hybrid Tables and Text
- Models Can Model, But Can't Bind: Structured Grounding in Text-to-Optimization
- MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference
- FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate
- Detailed balance in large language model-driven agents
- Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks
- SymStep: Symbolic Step Verification for Logical Reasoning
- TGMS: An Agent-Native Bi-Temporal Graph Management System
- MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities
- The Cartesian Cut in Agentic AI
- Policy-Conditioned Policies for Multi-Agent Task Solving
- AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
- Synthesizing Procedural Memory: Challenges and Architectures in Automated Workflow Generation
- External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
- Can Large Reasoning Models Improve Accuracy on Mathematical Tasks Using Flawed Thinking?
- Constructive Circuit Amplification: Improving Math Reasoning in LLMs via Targeted Sub-Network Updates
- BRAID: Bounded Reasoning for Autonomous Inference and Decisions
- Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision
- CodeMem: Architecting Reproducible Agents via Dynamic MCP and Procedural Memory
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- Intention Chain-of-Thought Prompting with Dynamic Routing for Code Generation
- Sharing State Between Prompts and Programs
- Error-Driven Prompt Optimization for Arithmetic Reasoning
- Revisiting the Reliability of Language Models in Instruction-Following
- PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
- Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
- JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
- TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models
- MARINE: Theoretical Optimization and Design for Multi-Agent Recursive IN-context Enhancement
- Auto-SPT: Automating Semantic Preserving Transformations for Code
- Thinking with Programming Vision: Towards a Unified View for Thinking with Images
- On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
- When Do Symbolic Solvers Enhance Reasoning in Large Language Models?
- CoRT: Code-integrated Reasoning within Thinking
- Generating Verifiable Chain of Thoughts from Exection-Traces
- ORION: Teaching Language Models to Reason Efficiently in the Language of Thought
- Asking like Socrates: Socrates helps VLMs understand remote sensing images
- On the Limits of Innate Planning in Large Language Models
- MedRule-KG: A Knowledge-Graph--Steered Scaffold for Reliable Mathematical and Biomedical Reasoning
- JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation
- Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
- GPS: General Per-Sample Prompter
- From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models
- Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
- Generative Caching for Structurally Similar Prompts and Responses
- Efficient Thought Space Exploration Through Strategic Intervention
- From Efficiency to Adaptivity: A Deeper Look at Adaptive Reasoning in Large Language Models
- | \circlearrowright \boxedBUS |: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
- Computational Blueprints: Generating Isomorphic Mathematics Problems with Large Language Models
- PCRLLM: Proof-Carrying Reasoning with Large Language Models under Stepwise Logical Constraints
- Better Datasets Start From RefineLab: Automatic Optimization for High-Quality Dataset Refinement
- RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables
- TabDSR: Decompose, Sanitize, and Reason for Complex Numerical Reasoning in Tabular Data
- Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
- Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective
- FlashEVA: Accelerating LLM inference via Efficient Attention
- ORGEval: Graph-Theoretic Evaluation of LLMs in Optimization Modeling
- TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models
- Mergeable Model-Side Aggregation States for Long-Context Language Models
- PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning
- Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?
- The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability
- GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring
- Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
- MedRule-KG: A Knowledge-Graph--Steered Scaffold for Mathematical Reasoning with a Lightweight Verifier
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
- Improving Human Verification of LLM Reasoning through Interactive Explanation Interfaces
- StreetMath: Study of LLMs' Approximation Behaviors
- Once Upon an Input: Reasoning via Per-Instance Program Synthesis
- Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
- Teaching Language Models to Reason with Tools
- Think Parallax: Solving Multi-Hop Problems via Multi-View Knowledge-Graph-Based Retrieval-Augmented Generation
- Code-enabled language models can outperform reasoning models on diverse tasks
- DelvePO: Direction-Guided Self-Evolving Framework for Flexible Prompt Optimization
- OPTAGENT: Optimizing Multi-Agent LLM Interactions Through Verbal Reinforcement Learning for Enhanced Reasoning
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models
- Investigating the Impact of Rationales for LLMs on Natural Language Understanding
- Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation
- Programmatic Representation Learning with Language Models
- MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning
- Where to Search: Measure the Prior-Structured Search Space of LLM Agents
- RECODE: Reasoning Through Code Generation for Visual Question Answering
- CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning
- Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative Retrieval
- Schema for In-Context Learning
- PricingLogic: Evaluating LLMs Reasoning on Complex Tourism Pricing Tasks
- LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
- Demystifying Reinforcement Learning in Agentic Reasoning
- ParaCook: On Time-Efficient Planning for Multi-Agent Systems
- Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
- Limits of Emergent Reasoning of Large Language Models in Agentic Frameworks for Deterministic Games
- FOR-Prompting: From Objection to Revision via an Asymmetric Prompting Protocol
- ARM2: Adaptive Reasoning Model with Vision Understanding and Executable Code
- GCPO: When Contrast Fails, Go Gold
- ToolLibGen: Scalable Automatic Tool Creation and Aggregation for LLM Reasoning
- BLAZER: Bootstrapping LLM-based Manipulation Agents with Zero-Shot Data Generation
- Revisiting the Uniform Information Density Hypothesis in LLM Reasoning Traces
- Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
- Adaptive Tool Generation with Models as Tools and Reinforcement Learning
- Expanding the Action Space of LLMs to Reason Beyond Language
- TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
- When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
- AlphaApollo: A System for Deep Agentic Reasoning
- Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
- MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information
- One More Question is Enough, Expert Question Decomposition (EQD) Model for Domain Quantitative Reasoning
- GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
- Lateral Tree-of-Thoughts Surpasses ToT by Incorporating Logically-Consistent, Low-Utility Candidates
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- TUMIX: Multi-Agent Test-Time Scaling with Tool-Use Mixture
- Plan before Solving: Problem-Aware Strategy Routing for Mathematical Reasoning with LLMs
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- Beyond Magic Words: Sharpness-Aware Prompt Evolving for Robust Large Language Models with TARE
- MedCritical: Enhancing Medical Reasoning in Small Language Models via Self-Collaborative Correction
- Decoupling Reasoning and Perception: An LLM-LMM Framework for Faithful Visual Reasoning
- Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning
- Reinforcement Learning-Guided Chain-of-Draft for Token-Efficient Code Generation
- Teaching Transformers to Solve Combinatorial Problems through Efficient Trial & Error
- ToMPO: Training LLM Strategic Decision Making from a Multi-Agent Perspective
- What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
- Distilling Answer Set Programming Theories from Large Language Models
- Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
- Solving Math Word Problems Using Estimation Verification and Equation Generation
- Understanding Benchmark Language Under Weakened Formal Semantics
- From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
- THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
- KoSEL: Knowledge subgraph enhanced large language model for medical question answering
- Improving Table Understanding with LLMs and Entity-Oriented Search
- Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
- Planning for Success: Exploring LLM Long-term Planning Capabilities in Table Understanding
- TORSO: Template-Oriented Reasoning Towards General Tasks
- How well can LLMs provide planning feedback in grounded environments?
- Accelerating Reinforcement Learning Algorithms Convergence using Pre-trained Large Language Models as Tutors With Advice Reusing
- CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning
- Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
- Rethinking the Chain-of-Thought: The Roles of In-Context Learning and Pre-trained Priors
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
- TableZoomer: A Collaborative Agent Framework for Large-scale Table Question Answering
- PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation
- CyberSleuth: Autonomous Blue-Team LLM Agent for Web Attack Forensics
- MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
- Thinking Before You Speak: A Proactive Test-time Scaling Approach
- MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning Chains
- Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
- XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
- Neuro-Symbolic Artificial Intelligence: Towards Improving the Reasoning Abilities of Large Language Models
- UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning
- Tabularis Formatus: Predictive Formatting for Tables
- Mathematical Computation and Reasoning Errors by Large Language Models
- MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
- NeuroSync: Intent-Aware Code-Based Problem Solving via Direct LLM Understanding Modification
- CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
- VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
- Text-to-SQL Task-oriented Dialogue Ontology Construction
- Failures Are the Stepping Stones to Success: Enhancing Few-Shot In-Context Learning by Leveraging Negative Samples
- A Toolbox, Not a Hammer -- Multi-TAG: Scaling Math Reasoning with Multi-Tool Aggregation
- Decoupling Knowledge and Reasoning in LLMs: An Exploration Using Cognitive Dual-System Theory
- Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
- SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models
- KiC: Keyword-inspired Cascade for Cost-Efficient Text Generation with LLMs
- Thinking Isn't an Illusion: Overcoming the Limitations of Reasoning Models via Tool Augmentations
- Manimator: Transforming Research Papers into Visual Explanations
- Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models
- A Survey of Deep Learning for Geometry Problem Solving
- Improved LLM Agents for Financial Document Question Answering
- Sample Efficient Demonstration Selection for In-Context Learning
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- Introspection of Thought Helps AI Agents
- From Language to Logic: A Bi-Level Framework for Structured Reasoning
- TableReasoner: Advancing Table Reasoning Framework with Large Language Models
- CRISP: Complex Reasoning with Interpretable Step-based Plans
- Agentic-R1: Distilled Dual-Strategy Reasoning
- AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction
- VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots
- Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?
- Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
- M3MAD-Bench: Multi-Dimensional Evaluation of Multi-Agent Debate Across Domains and Modalities
- Efficient Post-Training Refinement of Latent Reasoning in Large Language Models
- Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
- Learning-to-Context Slope: Evaluating In-Context Learning Effectiveness Beyond Performance Illusions
- Format-Adapter: Improving Reasoning Capability of LLMs by Adapting Suitable Format
- AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
- TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
- Distilling Tool Knowledge into Language Models via Back-Translated Traces
- Programming by Backprop: LLMs Acquire Reusable Algorithmic Abstractions During Code Training
- SWE-SQL: Illuminating LLM Pathways to Solve User SQL Issues in Real-World Applications
- Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action
- Chain of Methodologies: Scaling Test Time Computation without Training
- Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
- Structured Program Synthesis using LLMs: Results and Insights from the IPARC Challenge
- Graph of Verification: Structured Verification of LLM Reasoning with Directed Acyclic Graphs
- Efficient LLM Collaboration via Planning
- A Survey of Foundation Models for IoT: Taxonomy and Criteria-Based Analysis
- PRO-V-R1: Reasoning Enhanced Programming Agent for RTL Verification
- ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
- No Universal Prompt: Unifying Reasoning through Adaptive Prompting for Temporal Table Reasoning
- Table-r1: Self-supervised and Reinforcement Learning for Program-based Table Reasoning in Small Language Models
- FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging
- LLM-Symbolic Integration for Robust Temporal Tabular Reasoning
- Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models
- Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs
- Advancing Tool-Augmented Large Language Models via Meta-Verification and Reflection Learning
- Exchange of Perspective Prompting Enhances Reasoning in Large Language Models
- Verbalized Confidence Triggers Self-Verification: Emergent Behavior Without Explicit Reasoning Supervision
- Computational Thinking Reasoning in Large Language Models
- BitBypass: A New Direction in Jailbreaking Aligned Large Language Models with Bitstream Camouflage
- Decompose, Plan in Parallel, and Merge: A Novel Paradigm for Large Language Models based Planning with Multiple Constraints
- Read it in Two Steps: Translating Extremely Low-Resource Languages with Code-Augmented Grammar Books
- The Road to Generalizable Neuro-Symbolic Learning Should be Paved with Foundation Models
- Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research
- Semi-structured LLM Reasoners Can Be Rigorously Audited
- Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
- Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey
- Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
- Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation
- Fortune: Formula-Driven Reinforcement Learning for Symbolic Table Reasoning in Language Models
- InfiMed: Low-Resource Medical MLLMs with Advancing Understanding and Reasoning
- Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures
- Are Reasoning Models More Prone to Hallucination?
- Scalable, Symbiotic, AI and Non-AI Agent Based Parallel Discrete Event Simulations
- Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
- RefTool: Enhancing Model Reasoning with Reference-Guided Tool Creation
- R1-Code-Interpreter: LLMs Reason with Code via Supervised and Multi-stage Reinforcement Learning
- DecisionFlow: Advancing Large Language Model as Principled Decision Maker
- RRO: LLM Agent Optimization Through Rising Reward Trajectories
- Automatic Transmission for LLM Tiers: Optimizing Cost and Accuracy in Large Language Models
- MIRROR: Multi-agent Intra- and Inter-Reflection for Optimized Reasoning in Tool Learning
- HyperTree Planning: Enhancing LLM Reasoning via Hierarchical Thinking
- VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
- Program of Equations Thoughts to Solve Algebra Word Problems
- Position: Foundation Models for Tabular Data within Systemic Contexts Need Grounding
- Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers
- Interleaved Reasoning for Large Language Models via Reinforcement Learning
- Recursive Decomposition with Dependencies for Generic Divide-and-Conquer Reasoning
- Route to Reason: Adaptive Routing for LLM and Reasoning Strategy Selection
- Hierarchical Retrieval with Evidence Curation for Open-Domain Financial Question Answering on Standardized Documents
- Weaver: Interweaving SQL and LLM for Table Reasoning
- LiteCUA: Computer as MCP Server for Computer-Use Agent on AIOS
- UNJOIN: Enhancing Multi-Table Text-to-SQL Generation via Schema Simplification
- Reliable Natural Language Understanding with Large Language Models and Answer Set Programming
- Training with Pseudo-Code for Instruction Following
- InfoDet: A Dataset for Infographic Element Detection
- Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
- VLM-R3: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
- Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
- Logic-of-Thought: Empowering Large Language Models with Logic Programs for Solving Puzzles in Natural Language
- Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
- PRL: Prompts from Reinforcement Learning
- Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey
- From Reasoning to Code: GRPO Optimization for Underrepresented Languages
- Interpretable Traces, Unexpected Outcomes: Investigating the Disconnect in Trace-Based Knowledge Distillation
- DSMentor: Enhancing Data Science Agents with Curriculum Learning and Online Knowledge Accumulation
- RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
- Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
- AutoMathKG: The automated mathematical knowledge graph based on LLM and vector database
- Towards Functional Correctness of Large Code Models with Selective Generation
- MARGE: Improving Math Reasoning for LLMs with Guided Exploration
- Do Code LLMs Do Static Analysis?
- ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without Training
- GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents
- Ranked Voting based Self-Consistency of Large Language Models
- Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates
- Agentic Bayesian Optimization through Surrogate-Augmented Autoresearch
- Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving
- Evolutionary thoughts: integration of large language models and evolutionary algorithms
- Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages
- Brain-Inspired Graph Multi-Agent Systems for LLM Reasoning
- UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
- How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use
- Pencil Puzzle Bench: A Benchmark for Multi-Step Verifiable Reasoning
- Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks
- Symbolic-Neural Soft-Logic Reasoning: Towards Robust and Verifiable Thinking Chains via Cooperative Evolution
- DEL: Digit Entropy Loss for Numerical Learning of Large Language Models
- Agentic Test-Time Scaling for WebAgents
- Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games
- RV-Syn: Rational and Verifiable Mathematical Reasoning Data Synthesis based on Structured Function Library
- PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement
- HyperTool: Beyond Step-Wise Tool Calls for Tool-Augmented Agents
- The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge?
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
- VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing
- VeRA: Verified Reasoning Data Augmentation at Scale
- Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics
- AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
- Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models
- A Self-Improving Coding Agent
- a1: Steep Test-time Scaling Law via Environment Augmented Generation
- CoT-RAG: Integrating Chain of Thought and Retrieval-Augmented Generation to Enhance Reasoning in Large Language Models
- ToolRL: Reward is All Tool Learning Needs
- The Hitchhiker's Guide to Program Analysis, Part II: Deep Thoughts by LLMs
- Could Thinking Multilingually Empower LLM Reasoning?
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- Efficient Reasoning Models: A Survey
- Mixture-of-RAG: Integrating Text and Tables with Large Language Models
- Syzygy of Thoughts: Improving LLM CoT with the Minimal Free Resolution
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- A Desideratum for Conversational Agents: Capabilities, Challenges, and Future Directions
- ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
Discussions
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) [hn, 136 points, 36 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) https://arxiv.org/abs/2211.12588 [comments] [65 points] [bsky, 1 points, 0 comments]
- https://bsky.app/profile/buzzing.cc.web.brid.gy/post/3m6vzzdd2pnk2 [bsky, 0 points, 0 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) #HackerNews https://arxiv.org/abs/2211.12588 [bsky, 0 points, 0 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) https://arxiv.org/abs/2211.12588 https://news.ycombinator.com/item?id=46099108 [bsky, 0 points, 0 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) https://arxiv.org/abs/2211.12588 (https://news.ycombinator.com/item?id=46099108) [bsky, 0 points, 0 comments]
- ⚡ Hackernews Top story: Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) [bsky, 0 points, 0 comments]
- https://arxiv.org/abs/2211.12588 数値推論タスクにおいて、言語モデルに段階的な推論を実行させる研究が進んでいます。 Chain-of-Thoughts prompting (CoT)は最先端の手法ですが、推論と計算を両方行います。 本論文では、推論過程をプログラムとして表現するProgram of Thoughts (PoT)を提案し、計算を外部コンピュータに委ねま [bsky, 0 points, 0 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) https:// arxiv.org/abs/2211.12588 # arxiv [mastodon, 0 points, 0 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% https://arxiv.org/abs/2211.12588 [bsky, 0 points, 0 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) https://arxiv.org/abs/2211.12588 (https://news.ycombinator.com/item?id=46099108) [bsky, 0 points, 0 comments]
Related