Faithful Chain-of-Thought Reasoning
2023/01/31 by Qing Lyu, Lyu, Qing, Shreya Havaldar +13 · 1 voice · 96 citations
Computer Science · #Advanced Graph Neural Networks #Natural Language Processing Techniques #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2301.13379
arxiv published 2023/01/31 · arxiv updated 2023/09/20
Abstract
While Chain-of-Thought (CoT) prompting boosts Language Models' (LM) performance on a gamut of complex reasoning tasks, the generated reasoning chain does not necessarily reflect how the model arrives at the answer (aka. faithfulness). We propose Faithful CoT, a reasoning framework involving two stages: Translation (Natural Language query → symbolic reasoning chain) and Problem Solving (reasoning chain → answer), using an LM and a deterministic solver respectively. This guarantees that the reasoning chain provides a faithful explanation of the final answer. Aside from interpretability, Faithful CoT also improves empirical performance: it outperforms standard CoT on 9 of 10 benchmarks from 4 diverse domains, with a relative accuracy gain of 6.3% on Math Word Problems (MWP), 3.4% on Planning, 5.5% on Multi-hop Question Answering (QA), and 21.4% on Relational Inference. Furthermore, with GPT-4 and Codex, it sets the new state-of-the-art few-shot performance on 7 datasets (with 95.0+ accuracy on 6 of them), showing a strong synergy between faithfulness and accuracy.
Cited by
- SymStep: Symbolic Step Verification for Logical Reasoning
- Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming
- Semantic Deception: When Reasoning Models Can't Compute an Addition
- COIVis: Eye-tracking-based Visual Exploration of Concept Learning in MOOC Videos
- Training and Evaluation of Guideline-Based Medical Reasoning in LLMs
- See, Think, Learn: A Self-Taught Multimodal Reasoner
- More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering
- ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
- Cognitive Inception: Agentic Reasoning against Visual Deceptions by Injecting Skepticism
- From Efficiency to Adaptivity: A Deeper Look at Adaptive Reasoning in Large Language Models
- Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs
- Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
- MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
- Are Language Models Efficient Reasoners? A Perspective from Logic Programming
- GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
- The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
- SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Towards Transparent Reasoning: What Drives Faithfulness in Large Language Models?
- From Prompt Optimization to Multi-Dimensional Credibility Evaluation: Enhancing Trustworthiness of Chinese LLM-Generated Liver MRI Reports
- Once Upon an Input: Reasoning via Per-Instance Program Synthesis
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- Knowing the Facts but Choosing the Shortcut: Understanding How Large Language Models Compare Entities
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- PricingLogic: Evaluating LLMs Reasoning on Complex Tourism Pricing Tasks
- Verifying Chain-of-Thought Reasoning via Its Computational Graph
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
- Lateral Tree-of-Thoughts Surpasses ToT by Incorporating Logically-Consistent, Low-Utility Candidates
- DecepChain: Inducing Deceptive Reasoning in Large Language Models
- Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA
- From Faithfulness to Correctness: Generative Reward Models that Think Critically
- Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models
- Why Chain of Thought Fails in Clinical Text Understanding
- Vision Language Models Cannot Plan, but Can They Formalize?
- LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- Do Activation Verbalization Methods Convey Privileged Information?
- Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
- Unbiased Reasoning for Knowledge-Intensive Tasks in Large Language Models via Conditional Front-Door Adjustment
- Intermediate Languages Matter: Formal Languages and LLMs affect Neurosymbolic Reasoning
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
- Do Cognitively Interpretable Reasoning Traces Improve LLM Performance?
- AI reasoning effort predicts human decision time in content moderation
- Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
- Stands to Reason: Investigating the Effect of Reasoning on Idiomaticity Detection
- LLM Empowered Prototype Learning for Zero and Few-Shot Tasks on Tabular Data
- A Comparative Study of Neurosymbolic AI Approaches to Interpretable Logical Reasoning
- Thinking Machines: Mathematical Reasoning in the Age of LLMs
- Speaking in Words, Thinking in Logic: A Dual-Process Framework in QA Systems
- From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation
- Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization
- Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
- VERGE: Formal Refinement and Guidance Engine for Verifiable LLM Reasoning
- Probabilistic Soundness Guarantees in LLM Reasoning Chains
- Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Models
- KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
- Deep Hidden Cognition Facilitates Reliable Chain-of-Thought Reasoning
- Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
- Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
- Fast ECoT: Efficient Embodied Chain-of-Thought via Thoughts Reuse
- Reasoning Models Don't Always Say What They Think
- Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs
- When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
- AutoPatch: Multi-Agent Framework for Patching Real-World CVE Vulnerabilities
- Exchange of Perspective Prompting Enhances Reasoning in Large Language Models
- FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
- The Road to Generalizable Neuro-Symbolic Learning Should be Paved with Foundation Models
- ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generation
- Semi-structured LLM Reasoners Can Be Rigorously Audited
- CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
- Counterfactual Simulatability of LLM Explanations for Generation Tasks
- ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning
- CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
- Do Large Language Models Excel in Complex Logical Reasoning with Formal Language?
- Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
- Are LLMs Better Formalizers than Solvers on Complex Problems?
- CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process
- Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
- Reliable Post-Retrieval Assembly for Agent Memory: Separating Evidence Extraction from Policy Execution
- Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation
- Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory
- Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
- Interpretability Can Be Actionable
- Reasoning Models Will Sometimes Lie About Their Reasoning
- Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics
- The Rise of Small Language Models in Healthcare: A Comprehensive Survey
- AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models
- Reflexive Prompt Engineering: A Framework for Responsible Prompt Engineering and Interaction Design
- Probabilistic Stability Guarantees for Feature Attributions
- CoT-RAG: Integrating Chain of Thought and Retrieval-Augmented Generation to Enhance Reasoning in Large Language Models
- PEIRCE: Unifying Material and Formal Reasoning via LLM-Driven Neuro-Symbolic Refinement
- Inherent and emergent liability issues in LLM-based agentic systems: a principal-agent perspective
- Bonsai: Interpretable Tree-Adaptive Grounded Reasoning
Discussions
Related