Large Language Models Imitate Logical Reasoning, but at what Cost?
2025/09/16 by Lachlan McGinness, McGinness, Lachlan, Peter Baumgartner +1 · 1 citation
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Semantic Web and Ontologies
paper · pdf · doi:10.48550/arxiv.2509.12645
Abstract
We present a longitudinal study which evaluates the reasoning capability of frontier Large Language Models over an eighteen month period. We measured the accuracy of three leading models from December 2023, September 2024 and June 2025 on true or false questions from the PrOntoQA dataset and their faithfulness to reasoning strategies provided through in-context learning. The improvement in performance from 2023 to 2024 can be attributed to hidden Chain of Thought prompting. The introduction of thinking models allowed for significant improvement in model performance between 2024 and 2025. We then present a neuro-symbolic architecture which uses LLMs of less than 15 billion parameters to translate the problems into a standardised form. We then parse the standardised forms of the problems into a program to be solved by Z3, an SMT solver, to determine the satisfiability of the query. We report the number of prompt and completion tokens as well as the computational cost in FLOPs for open source models. The neuro-symbolic approach significantly reduces the computational cost while maintaining near perfect performance. The common approximation that the number of inference FLOPs is double the product of the active parameters and total tokens was accurate within 10% for all experiments.
Citations
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Efficient Inference for Large Reasoning Models: A Survey
- Gemma 3 Technical Report
- Logical Reasoning in Large Language Models: A Survey
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Phi-4 Technical Report
- Enhancing Reasoning Capabilities of LLMs via Principled Synthetic Logic Corpus
- The Llama 3 Herd of Models
- Steamroller Problems: An Evaluation of LLM Reasoning Capability with Automated Theorem Prover Strategies
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Let's Think Dot by Dot: Hidden Computation in Transformer Language Models
- Many-Shot In-Context Learning
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Power Hungry Processing: Watts Driving the Cost of AI Deployment?
- Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond
- Faith and Fate: Limits of Transformers on Compositionality
- On the Planning Abilities of Large Language Models : A Critical Investigation
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- GPT-4 Technical Report
- LAMBADA: Backward Chaining for Automated Reasoning in Natural Language
- Towards Reasoning in Large Language Models: A Survey
- Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- Training Compute-Optimal Large Language Models
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- On the Dangers of Stochastic Parrots
- Language Models are Few-Shot Learners
- Scaling Laws for Neural Language Models
- Root Mean Square Layer Normalization
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
- Attention Is All You Need
- Using the Output Embedding to Improve Language Models
- Gaussian Error Linear Units (GELUs)
- Probable Inference, the Law of Succession, and Statistical Inference
Cited by
Related