Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
2022/11/22 by Wenhu Chen, Chen, Wenhu, Xueguang Ma +5 · 11 voices · 166 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Online Learning and Analytics #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2211.12588
openalex publication_date 2022/11/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Recently, there has been significant progress in teaching language models to perform step-by-step reasoning to solve complex numerical reasoning tasks. Chain-of-thoughts prompting (CoT) is by far the state-of-art method for these tasks. CoT uses language models to perform both reasoning and computation in the multi-step `thought' process. To disentangle computation from reasoning, we propose `Program of Thoughts' (PoT), which uses language models (mainly Codex) to express the reasoning process as a program. The computation is relegated to an external computer, which executes the generated programs to derive the answer. We evaluate PoT on five math word problem datasets (GSM, AQuA, SVAMP, TabMWP, MultiArith) and three financial-QA datasets (FinQA, ConvFinQA, TATQA) for both few-shot and zero-shot setups. Under both few-shot and zero-shot settings, PoT can show an average performance gain over CoT by around 12% across all the evaluated datasets. By combining PoT with self-consistency decoding, we can achieve SoTA performance on all math problem datasets and near-SoTA performance on financial datasets. All of our data and code are released in Github https://github.com/wenhuchen/Program-of-Thoughts
Cited by
- Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models
- NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
- TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
- Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
- PAGE-RAG: Evidence-Grounded Adaptive Graph Retrieval for Long-Document Question Answering
- Gold-Guided Programmatic Distillation for Financial Reasoning over Hybrid Tables and Text
- Models Can Model, But Can't Bind: Structured Grounding in Text-to-Optimization
- MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference
- FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate
- Detailed balance in large language model-driven agents
- Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks
- SymStep: Symbolic Step Verification for Logical Reasoning
- TGMS: An Agent-Native Bi-Temporal Graph Management System
- MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities
- The Cartesian Cut in Agentic AI
- Policy-Conditioned Policies for Multi-Agent Task Solving
- AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
- Synthesizing Procedural Memory: Challenges and Architectures in Automated Workflow Generation
- External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
- Can Large Reasoning Models Improve Accuracy on Mathematical Tasks Using Flawed Thinking?
- Constructive Circuit Amplification: Improving Math Reasoning in LLMs via Targeted Sub-Network Updates
- BRAID: Bounded Reasoning for Autonomous Inference and Decisions
- Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision
- CodeMem: Architecting Reproducible Agents via Dynamic MCP and Procedural Memory
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- Intention Chain-of-Thought Prompting with Dynamic Routing for Code Generation
- Sharing State Between Prompts and Programs
- Error-Driven Prompt Optimization for Arithmetic Reasoning
- Revisiting the Reliability of Language Models in Instruction-Following
- PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
- Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
- JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
- TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models
- MARINE: Theoretical Optimization and Design for Multi-Agent Recursive IN-context Enhancement
- Auto-SPT: Automating Semantic Preserving Transformations for Code
- Thinking with Programming Vision: Towards a Unified View for Thinking with Images
- On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
- When Do Symbolic Solvers Enhance Reasoning in Large Language Models?
- Generating Verifiable Chain of Thoughts from Exection-Traces
- ORION: Teaching Language Models to Reason Efficiently in the Language of Thought
- Asking like Socrates: Socrates helps VLMs understand remote sensing images
- On the Limits of Innate Planning in Large Language Models
- MedRule-KG: A Knowledge-Graph--Steered Scaffold for Reliable Mathematical and Biomedical Reasoning
- JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation
- Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
- GPS: General Per-Sample Prompter
- From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models
- Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
- Generative Caching for Structurally Similar Prompts and Responses
- Efficient Thought Space Exploration Through Strategic Intervention
- From Efficiency to Adaptivity: A Deeper Look at Adaptive Reasoning in Large Language Models
- | \circlearrowright \boxedBUS |: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
- Computational Blueprints: Generating Isomorphic Mathematics Problems with Large Language Models
- PCRLLM: Proof-Carrying Reasoning with Large Language Models under Stepwise Logical Constraints
- Better Datasets Start From RefineLab: Automatic Optimization for High-Quality Dataset Refinement
- RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables
- TabDSR: Decompose, Sanitize, and Reason for Complex Numerical Reasoning in Tabular Data
- Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
- Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective
- FlashEVA: Accelerating LLM inference via Efficient Attention
- ORGEval: Graph-Theoretic Evaluation of LLMs in Optimization Modeling
- TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models
- Mergeable Model-Side Aggregation States for Long-Context Language Models
- PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning
- Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?
- The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability
- GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring
- Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
- MedRule-KG: A Knowledge-Graph--Steered Scaffold for Mathematical Reasoning with a Lightweight Verifier
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
- Improving Human Verification of LLM Reasoning through Interactive Explanation Interfaces
- StreetMath: Study of LLMs' Approximation Behaviors
- Once Upon an Input: Reasoning via Per-Instance Program Synthesis
- Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
- Teaching Language Models to Reason with Tools
- Think Parallax: Solving Multi-Hop Problems via Multi-View Knowledge-Graph-Based Retrieval-Augmented Generation
- Code-enabled language models can outperform reasoning models on diverse tasks
- DelvePO: Direction-Guided Self-Evolving Framework for Flexible Prompt Optimization
- OPTAGENT: Optimizing Multi-Agent LLM Interactions Through Verbal Reinforcement Learning for Enhanced Reasoning
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models
- Investigating the Impact of Rationales for LLMs on Natural Language Understanding
- Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation
- Programmatic Representation Learning with Language Models
- MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning
- Where to Search: Measure the Prior-Structured Search Space of LLM Agents
- RECODE: Reasoning Through Code Generation for Visual Question Answering
- CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning
- Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative Retrieval
- Schema for In-Context Learning
- PricingLogic: Evaluating LLMs Reasoning on Complex Tourism Pricing Tasks
- LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
- Demystifying Reinforcement Learning in Agentic Reasoning
- ParaCook: On Time-Efficient Planning for Multi-Agent Systems
- Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
- Limits of Emergent Reasoning of Large Language Models in Agentic Frameworks for Deterministic Games
- FOR-Prompting: From Objection to Revision via an Asymmetric Prompting Protocol
- ARM2: Adaptive Reasoning Model with Vision Understanding and Executable Code
- GCPO: When Contrast Fails, Go Gold
- ToolLibGen: Scalable Automatic Tool Creation and Aggregation for LLM Reasoning
- BLAZER: Bootstrapping LLM-based Manipulation Agents with Zero-Shot Data Generation
- Revisiting the Uniform Information Density Hypothesis in LLM Reasoning Traces
- Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
- Adaptive Tool Generation with Models as Tools and Reinforcement Learning
- Expanding the Action Space of LLMs to Reason Beyond Language
- TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
- When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
- AlphaApollo: Orchestrating Foundation Models and Professional Tools into a Self-Evolving System for Deep Agentic Reasoning
- Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
- MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information
- One More Question is Enough, Expert Question Decomposition (EQD) Model for Domain Quantitative Reasoning
- GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
- Lateral Tree-of-Thoughts Surpasses ToT by Incorporating Logically-Consistent, Low-Utility Candidates
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- TUMIX: Multi-Agent Test-Time Scaling with Tool-Use Mixture
- Plan before Solving: Problem-Aware Strategy Routing for Mathematical Reasoning with LLMs
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- Beyond Magic Words: Sharpness-Aware Prompt Evolving for Robust Large Language Models with TARE
- MedCritical: Enhancing Medical Reasoning in Small Language Models via Self-Collaborative Correction
- Decoupling Reasoning and Perception: An LLM-LMM Framework for Faithful Visual Reasoning
- Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning
- Reinforcement Learning-Guided Chain-of-Draft for Token-Efficient Code Generation
- Teaching Transformers to Solve Combinatorial Problems through Efficient Trial & Error
- ToMPO: Training LLM Strategic Decision Making from a Multi-Agent Perspective
- What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
- Distilling Answer Set Programming Theories from Large Language Models
- Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
- Solving Math Word Problems Using Estimation Verification and Equation Generation
- Understanding Benchmark Language Under Weakened Formal Semantics
- From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
- THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
- KoSEL: Knowledge subgraph enhanced large language model for medical question answering
- Improving Table Understanding with LLMs and Entity-Oriented Search
- Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
- Planning for Success: Exploring LLM Long-term Planning Capabilities in Table Understanding
- TORSO: Template-Oriented Reasoning Towards General Tasks
- How well can LLMs provide planning feedback in grounded environments?
- Accelerating Reinforcement Learning Algorithms Convergence using Pre-trained Large Language Models as Tutors With Advice Reusing
- CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning
- Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
- Rethinking the Chain-of-Thought: The Roles of In-Context Learning and Pre-trained Priors
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
- TableZoomer: A Collaborative Agent Framework for Large-scale Table Question Answering
- PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation
- CyberSleuth: Autonomous Blue-Team LLM Agent for Web Attack Forensics
- MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
- Thinking Before You Speak: A Proactive Test-time Scaling Approach
- MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning Chains
- Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
- XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
- Neuro-Symbolic Artificial Intelligence: Towards Improving the Reasoning Abilities of Large Language Models
- UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning
- Tabularis Formatus: Predictive Formatting for Tables
- Mathematical Computation and Reasoning Errors by Large Language Models
- MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
- NeuroSync: Intent-Aware Code-Based Problem Solving via Direct LLM Understanding Modification
- CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
- Text-to-SQL Task-oriented Dialogue Ontology Construction
- Failures Are the Stepping Stones to Success: Enhancing Few-Shot In-Context Learning by Leveraging Negative Samples
- A Toolbox, Not a Hammer -- Multi-TAG: Scaling Math Reasoning with Multi-Tool Aggregation
- Decoupling Knowledge and Reasoning in LLMs: An Exploration Using Cognitive Dual-System Theory
- SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models
Discussions
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) [hn, 136 points, 36 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) https://arxiv.org/abs/2211.12588 [comments] [65 points] [bsky, 1 points, 0 comments]
- https://bsky.app/profile/buzzing.cc.web.brid.gy/post/3m6vzzdd2pnk2 [bsky, 0 points, 0 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) #HackerNews https://arxiv.org/abs/2211.12588 [bsky, 0 points, 0 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) https://arxiv.org/abs/2211.12588 https://news.ycombinator.com/item?id=46099108 [bsky, 0 points, 0 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) https://arxiv.org/abs/2211.12588 (https://news.ycombinator.com/item?id=46099108) [bsky, 0 points, 0 comments]
- ⚡ Hackernews Top story: Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) [bsky, 0 points, 0 comments]
- https://arxiv.org/abs/2211.12588 数値推論タスクにおいて、言語モデルに段階的な推論を実行させる研究が進んでいます。 Chain-of-Thoughts prompting (CoT)は最先端の手法ですが、推論と計算を両方行います。 本論文では、推論過程をプログラムとして表現するProgram of Thoughts (PoT)を提案し、計算を外部コンピュータに委ねま [bsky, 0 points, 0 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) https:// arxiv.org/abs/2211.12588 # arxiv [mastodon, 0 points, 0 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% https://arxiv.org/abs/2211.12588 [bsky, 0 points, 0 comments]
- Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022) https://arxiv.org/abs/2211.12588 (https://news.ycombinator.com/item?id=46099108) [bsky, 0 points, 0 comments]
Related