GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
2024/10/07 by Iman Mirzadeh, Keivan Alizadeh, Mirzadeh, Iman +9 · 69 voices · 142 citations
Computer Science · #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2410.05229
Abstract
Recent advancements in Large Language Models (LLMs) have sparked interest in their formal reasoning capabilities, particularly in mathematics. The GSM8K benchmark is widely used to assess the mathematical reasoning of models on grade-school-level questions. While the performance of LLMs on GSM8K has significantly improved in recent years, it remains unclear whether their mathematical reasoning capabilities have genuinely advanced, raising questions about the reliability of the reported metrics. To address these concerns, we conduct a large-scale study on several SOTA open and closed models. To overcome the limitations of existing evaluations, we introduce GSM-Symbolic, an improved benchmark created from symbolic templates that allow for the generation of a diverse set of questions. GSM-Symbolic enables more controllable evaluations, providing key insights and more reliable metrics for measuring the reasoning capabilities of models.Our findings reveal that LLMs exhibit noticeable variance when responding to different instantiations of the same question. Specifically, the performance of all models declines when only the numerical values in the question are altered in the GSM-Symbolic benchmark. Furthermore, we investigate the fragility of mathematical reasoning in these models and show that their performance significantly deteriorates as the number of clauses in a question increases. We hypothesize that this decline is because current LLMs cannot perform genuine logical reasoning; they replicate reasoning steps from their training data. Adding a single clause that seems relevant to the question causes significant performance drops (up to 65%) across all state-of-the-art models, even though the clause doesn't contribute to the reasoning chain needed for the final answer. Overall, our work offers a more nuanced understanding of LLMs' capabilities and limitations in mathematical reasoning.
Cited by
- Semiotic logical hexagon theory for LLM logical reasoning
- Implicit Reasoning Steering via Concept Chaining
- ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
- Robust Reasoning Benchmark
- Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
- Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset Generator
- The Language of Security: How Prompt Syntax Shapes Secure Code Generation in Open LLMs
- Inductive Risk of AI Hype
- A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction
- Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
- Provable Scaling Laws of Feature Emergence from Learning Dynamics of Grokking
- Is In-Context Learning Learning?
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- Modeling Open-World Cognition as On-Demand Synthesis of Probabilistic Models
- Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
- Large Language Models and Emergence: A Complex Systems Perspective
- Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis
- On the Limits of LLM Reasoning: Evidence From Contamination, Translation, and Answer Modification in Multiple-Choice Benchmarks
- General Intelligence Requires Reward-based Pretraining
- LIMO: Less is More for Reasoning
- Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?
- Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation
- Can LLMs Write Correct TLA+ Specifications? Evaluating Natural-Language-to-TLA+ Generation
- Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
- Algorithmic Blindness in Large Language Models: A Calibration Study of Performance Prediction
- FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
- Concept Generalization in Humans and Large Language Models: Insights from the Number Game
- CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical Reasoning
- Constructive Circuit Amplification: Improving Math Reasoning in LLMs via Targeted Sub-Network Updates
- In-Context Algebra
- Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections
- Automated Penetration Testing with LLM Agents and Classical Planning
- Opportunities and Challenges in Harnessing Digital Technology for Effective Teaching and Learning
- Towards Language Model Guided TLA+ Proof Automation
- Neurosymbolic Information Extraction from Transactional Documents
- Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
- BEAVER: An Efficient Deterministic LLM Verifier
- EngChain: A Symbolic Benchmark for Verifiable Multi-Step Reasoning in Engineering
- WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
- Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
- A perceptual bias of AI Logical Argumentation Ability in Writing
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- Knowing What You Know Is Not Enough: Large Language Model Confidences Don't Align With Their Actions
- Structured Decomposition for LLM Reasoning: Cross-Domain Validation and Semantic Web Integration
- LM4Opt-RA: A Multi-Candidate LLM Framework with Structured Ranking for Automating Network Resource Allocation
- On the Notion that Language Models Reason
- Uncertainty-Guided Checkpoint Selection for Reinforcement Finetuning of Large Language Models
- The Ouroboros of Benchmarking: Reasoning Evaluation in an Era of Saturation
- MSCR: Exploring the Vulnerability of LLMs' Mathematical Reasoning Abilities Using Multi-Source Candidate Replacement
- Towards a Standard, Enterprise-Relevant Agentic AI Benchmark: Lessons from 5.5 billion tokens' worth of agentic AI evaluations
- Numerical Sensitivity and Robustness: Exploring the Flaws of Mathematical Reasoning in Large Language Models
- RIDE: Difficulty Evolving Perturbation with Item Response Theory for Mathematical Reasoning
- Batch Prompting Suppresses Overthinking Reasoning Under Constraint: How Batch Prompting Suppresses Overthinking in Reasoning Models
- Contamination Detection for VLMs using Multi-Modal Semantic Perturbation
- In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
- Neurosymbolic Deep Learning Semantics
- DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- Ariadne: A Controllable Framework for Probing and Extending VLM Reasoning Boundaries
- User Misconceptions of LLM-Based Conversational Programming Assistants
- Are Language Models Efficient Reasoners? A Perspective from Logic Programming
- Can we use automated approaches to measure the quality of online political discussion? How to (not) measure interactivity, diversity, rationality, and incivility in online comments to the news
- Transformers Can Learn Rules They've Never Seen: Proof of Computation Beyond Interpolation
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Human-Level Reasoning: A Comparative Study of Large Language Models on Logical and Abstract Reasoning
- SynthWorlds: Controlled Parallel Worlds for Disentangling Reasoning and Knowledge in Language Models
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
- StreetMath: Study of LLMs' Approximation Behaviors
- Modeling Hierarchical Thinking in Large Reasoning Models
- SEGA: A Stepwise Evolution Paradigm for Content-Aware Layout Generation with Design Prior
- Exploring Spiking Neural Networks for Binary Classification in Multivariate Time Series at the Edge
- What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
- Illusions of reflection: open-ended task reveals systematic failures in Large Language Models' reflective reasoning
- Adaptive Coopetition: Leveraging Coarse Verifier Signals for Resilient Multi-Agent LLM Reasoning
- Evaluating LLM Reasoning Beyond Correctness and CoT
- StreamingThinker: Large Language Models Can Think While Reading
- I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
- Interpretability Framework for LLMs in Undergraduate Calculus
- UniCode: A Framework for Generating High Quality Competitive Coding Problems
- Rethinking Evaluation in the Era of Time Series Foundation Models: (Un)known Information Leakage Challenges
- Interpreting the Latent Structure of Operator Precedence in Language Models
- MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
- Agent-Based Simulation of a Financial Market with Large Language Models
- Evaluating Language Models' Evaluations of Games
- Revisiting the UID Hypothesis in LLM Reasoning Traces
- MaP: A Unified Framework for Reliable Evaluation of Pre-training Dynamics
- RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows
- Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Models
- ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
- Efficient numeracy in language models through single-token number embeddings
- Iterative LLM-Based Generation and Refinement of Distracting Conditions in Math Word Problems
- h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
- RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
- Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
- Making Mathematical Reasoning Adaptive
- Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations
- MetaLogic: Robustness Evaluation of Text-to-Image Models via Logically Equivalent Prompts
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
- Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
- BloomAPR: A Bloom's Taxonomy-based Framework for Assessing the Capabilities of LLM-Powered APR Solutions
- From Ambiguity to Verdict: A Semiotic-Grounded Multi-Perspective Agent for LLM Logical Reasoning
- Evaluating Program Semantics Reasoning with Type Inference in System F
- Developing foundations for biomedical knowledgebases from literature using large language models – A systematic assessment
- Tracing Uncertainty in Language Model "Reasoning"
- From Deferral to Learning: Online In-Context Knowledge Distillation for LLM Cascades
- Psychometrically derived 60-question benchmarks: Substantial efficiencies and the possibility of human-AI comparisons
- IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
- GSM-Agent: Understanding Agentic Reasoning Using Controllable Environments
- Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
- Quantifying the Impact of Structured Output Format on Large Language Models through Causal Inference
- Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
- MARS: toward more efficient multi-agent collaboration for LLM reasoning
- Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
- Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For Perplexity
- EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- Robustness of Neurosymbolic Reasoners on First-Order Logic Problems
- Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs
- Toward Efficient Influence Function: Dropout as a Compression Tool
- Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery
- Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
- Investigating Student Interaction Patterns with Large Language Model-Powered Course Assistants in Computer Science Courses
- RL Fine-Tuning Heals OOD Forgetting in SFT
- MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations
- RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
- Self-Aligned Reward: Towards Effective and Efficient Reasoners
- CausalARC: Abstract Reasoning with Causal World Models
- Generative KI für TA
- Graph RAG as Human Choice Model: Building a Data-Driven Mobility Agent with Preference Chain
- CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs
- Robustness is Important: Limitations of LLMs for Data Fitting
- Even Heads Fix Odd Errors: Mechanistic Discovery and Surgical Repair in Transformer Attention
- Teaching LLMs to Think Mathematically: A Critical Study of Decision-Making via Optimization
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
- SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows
- Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
- ViPro-2: Unsupervised State Estimation via Integrated Dynamics for Guiding Video Prediction
- CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
- Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data
- Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities
- OpenAI o1 [wikipedia]
Discussions
- Understanding the Limitations of Mathematical Reasoning in LLMs [hn, 282 points, 266 comments]
- - "Reasoning"? Lets set aside how they don't even have a definition for this. But literally change some minor thing on the benchmarks like a number, and you see how these models completely fail. arxiv [bsky, 53 points, 1 comments]
- Can LLM-based systems "reason"? A recent paper arxiv.org/abs/2410.05229 looks at performance on word problems, and hypothesizes that "current LLMs are not capable of genuine logical reasoning." Let's [bsky, 34 points, 4 comments]
- Important, hype-puncturing paper.: "current LLMs are not capable of genuine logical reasoning; instead, they attempt to replicate the reasoning steps observed in their training data." AGI ain't around [bsky, 27 points, 2 comments]
- Mathematical proof that LLMs cannot reason, only act as pattern-matching machines, courtesy of Apple. (Explains why they (Apple) were first to jump back off the AI bandwagon) arxiv.org/pdf/2410.052 [bsky, 11 points, 0 comments]
- " LLMs struggle even when provided with multiple examples of the same question or examples containing similar irrelevant information. This suggests deeper issues in their reasoning processes that cann [bsky, 7 points, 2 comments]
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models arxiv.org/pdf/2410.05229 "current LLMs are not capable of genuine logical reasoning; instead, they attem [bsky, 6 points, 1 comments]
- current LLMs cannot perform genuine logical reasoning; they replicate reasoning steps from their training data. Adding a single clause that seems relevant to the question causes significant performanc [bsky, 5 points, 0 comments]
- LLMs still exhibit substantial limitations in logical reasoning. This means they are better at identifying familiar patterns but not necessarily improving in “thinking” or logical deduction. #AI #Pat [bsky, 4 points, 0 comments]
- arxiv.org/pdf/2410.05229 Apple researchers suggest artificial intelligence is still mostly an illusion. #AI #artificialintelligence [bsky, 4 points, 0 comments]
- White paper link to skip the paywall: arxiv.org/pdf/2410.05229 [bsky, 3 points, 0 comments]
- As this recent paper found - LLMs basically work like kids who did rote learning on 3 examples. They've got the samples down pat. But they immediately illogically riff once asked to do minute critical [bsky, 3 points, 2 comments]
- This pre-print paper doesn't tell us Luddites anything we didn't already know, but for your more AI-curious friends, please remind them that computers can't think, can't reason. arxiv.org/pdf/2410.05 [bsky, 3 points, 1 comments]
- “It may resemble sophisticated pattern matching more than true logical reasoning. We remind the reader that both GSM8K and GSM-Symbolic include relatively simple grade-school math questions, requiring [bsky, 3 points, 0 comments]
- Interesting feature of the Apple LLM reasoning paper. I always tell my students that exams include no irrelevant information, which gives you a clue as to the answer. LLM's have learnt this too well, [bsky, 3 points, 1 comments]
- GSM-Symbolic [lobsters, 2 points, 1 comments]
- "GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models" arxiv.org/pdf/2410.05229 [bsky, 2 points, 0 comments]
- Está probado (arxiv.org/pdf/2410.05229) que cuando un LLM recibe información desconocida o intrascendente sobre una tarea conocida, su rendimiento cae hasta en un 60%. Por otro lado, creo que otra pa [bsky, 2 points, 2 comments]
- arxiv.org/abs/2410.052... [bsky, 2 points, 0 comments]
- Check out this paper from Apple which throws cold water on the idea that LLMS can reason. arxiv.org/pdf/2410.05229 or this video which goes into detail on the paper: youtu.be/tTG_a0KPJAc?... I’m incl [bsky, 2 points, 0 comments]
- Extremely interesting new paper from AI researchers at Apple. arxiv.org/abs/2410.05229 [bsky, 2 points, 1 comments]
- And speaking of math olympiads: such human would also not commit these mistakes. arxiv.org/abs/2410.05229 [bsky, 2 points, 1 comments]
- Large language models (~AI) are not yet capable of true reasoning. Even slight changes in numerical values or problem complexity can cause significant performance drops. Not capable of logical reasoni [bsky, 2 points, 0 comments]
- #AMIA2025 #MLSky Direct link to the paper referenced: arxiv.org/abs/2410.05229 [bsky, 1 points, 0 comments]
- The paper on ARXIV. arxiv.org/pdf/2410.05229 [bsky, 1 points, 0 comments]
- Also: arxiv.org/abs/2410.05229 [bsky, 1 points, 1 comments]
- Game of AI: Winter Is Coming arxiv.org/abs/2410.05229 [bsky, 1 points, 0 comments]
- Lugemismaterjali sulle :) arxiv.org/pdf/2410.05229 LLMid on "statistilised masinad" ja nende "info" vajab alati põhjalikku järgikontrollimist :) Ma loodan, et sa seda seekord ka tegid enne oma väidet [bsky, 1 points, 1 comments]
- I'm not sure a statistical parrot (LLM based AI) will have much luck there. If a human had internalised all of wikipedia they would probably be a creative polymath with numerous Nobel prizes. Instead [bsky, 1 points, 1 comments]
- This presupposes that this growth will continue, but it will not. arxiv.org/abs/2410.05229 [bsky, 1 points, 1 comments]
- Worthwhile read for those enamored with AI arxiv.org/pdf/2410.05229 It just isn't, nor was it ever "I." Doesn't mean it can't be useful for sifting through data/text, identifying patterns & projecting [bsky, 1 points, 2 comments]
- Or this paper from late 2024, which looks at "reasoning" capabilities of probabilistic models and ultimately concludes they can't be accurate without formal reasoning capabilities. arxiv.org/abs/2410. [bsky, 1 points, 0 comments]
- Wirklich interessante Studie zu den Fähigkeiten logischen Schlussfolgerns von LLM: „GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models“ (arxiv.org/abs/2410. [bsky, 1 points, 1 comments]
- LLMs are not designed for logical reasoning and they are not good at it. arxiv.org/pdf/2410.05229 [bsky, 1 points, 1 comments]
- nice, thanks for sharing, i was thinking along of what apple did with their studies into reasoning: arxiv.org/pdf/2410.05229 [bsky, 1 points, 1 comments]
- Mooi dat BMGN zoveel expertise bij elkaar brengt! Er zijn top-experts die vinden dat de waarde van ChatGPT en andere LLM's wordt overschat. Het is ook business, bedoeld om ons te binden, terwijl zelf [bsky, 1 points, 1 comments]
- Understanding the Limitations of Mathematical Reasoning in LLMs: « Current LLMs cannot perform genuine logical reasoning; they replicate reasoning steps from their training data. » arxiv.org/abs/2410. [bsky, 1 points, 1 comments]
- Bad news for the #LLM #GenAI faithful - new study by Apple shows that LLM's do NOT show real logical reasoning in mathematics - they attempt to replicate the reasoning steps observed in their trainin [bsky, 1 points, 0 comments]
- The paper is focused on reasoning models and wants to determine whether they are capable of generalizable reasoning. This looks to be building on some of their previous work arxiv.org/abs/2410.05229 w [bsky, 0 points, 1 comments]
- Understanding the Limitations of Mathematical Reasoning in Large Language Models Recent advancements in Large Language Models (LLMs) have sparked interest in their formal reasoning capabilities, part [bsky, 0 points, 0 comments]
- Es gab doch erst Anfang Oktober die Studie von Apple Forschern, die gezeigt hat, dass es keine Anzeichen von „Reasoning“ in aktuellen LLMs gibt. Studie: arxiv.org/pdf/2410.05229 [bsky, 0 points, 0 comments]
- Apple: LLMs are bad at logic, reasoning, math https://arxiv.org/pdf/2410.05229 Google: so are our execs! https://www.theguardian.com/technology/2024/oct/15/google-buy-nuclear-power-ai-datacentres-kai [bsky, 0 points, 0 comments]
- The science is actually pretty clear, LLM’s do not reason and what they are doing does not resemble reasoning closely enough to develop in that direction. Gen AI is possible, LLM research may move us [bsky, 0 points, 1 comments]
- Good to see investigations into the assertions being made about LLMs. “We hypothesize that this decline is because current LLMs cannot perform genuine logical reasoning; they replicate reasoning step [bsky, 0 points, 0 comments]
- Apple: LLMs are bad at logic, reasoning, math arxiv.org/pdf/2410.05229 Google: so are our execs! www.theguardian.com/technology/2... [bsky, 0 points, 0 comments]
- Yeah but basically I'm extremely confident that the only thing that can be delegated is data labeling. The stuff that Golub mentions these models are better at is just wrong. As an example, here's a g [bsky, 0 points, 1 comments]
- 📌 📖 📑 arxiv.org/abs/2410.05229 [bsky, 0 points, 0 comments]
- I have little knowledge of LLM, my understanding is that most are 'predictions about the next word' why one would do maths testing? of course it is sensitive to input and order, that's how they are bu [bsky, 0 points, 0 comments]
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models. "Overall, we find that models tend to convert statements to operations without truly understanding their [bsky, 0 points, 0 comments]
- Apple research shows that LLMs have inherent limitations that make them unfit as agents. That’s too bad - LLMs will have their uses, but we’ll need a different breakthrough for actual intelligence. [bsky, 0 points, 0 comments]
- The mentioned paper: arxiv.org/pdf/2410.05229 [bsky, 0 points, 0 comments]
- Understanding the Limitations of Mathematical Reasoning in Large Language Models [bsky, 0 points, 0 comments]
- For the people in the back - LLM AI like GPT models cannot reason. It's literally not possible given the technology. Here's a study that lays it out neatly - arxiv.org/pdf/2410.05229 [bsky, 0 points, 0 comments]
- LLM-AI’s Limitations: arxiv.org/pdf/2410.05229 [bsky, 0 points, 0 comments]
- Can LLMs) truly reason? Or are they just sophisticated pattern matchers? Apple explores this key question through a large-scale study of both open-source like Llama, Phi, Gemma, and Mistral and leadin [bsky, 0 points, 0 comments]
- arxiv.org/pdf/2410.05229 #compsci #ai #chatgpt #deepseek #grok [bsky, 0 points, 0 comments]
- 🤔 arxiv.org/pdf/2410.05229 [bsky, 0 points, 1 comments]
- New paper by apple concludes there is no evidence of formal reasoning in LLM’s. arxiv.org/pdf/2410.05229 #AI #llm #reasoning #stochasticparrot [bsky, 0 points, 0 comments]
- This paper explained things well for me and you're pretty spot on. arxiv.org/abs/2410.05229 [bsky, 0 points, 0 comments]
- https://buff.ly/4hda1Wy [bsky, 0 points, 0 comments]
- Hi! Yeah I was a bit sarcastic 😅 the plot is from arxiv.org/abs/2410.05229 [bsky, 0 points, 0 comments]
- well, here are also a bunch of papers that show llms are better seen as statistical pattern matching 1. arxiv.org/pdf/2410.05229 2. arxiv.org/pdf/2307.02477 3. arxiv.org/pdf/2406.11012 4. www.pnas.org [bsky, 0 points, 1 comments]
- Are LLMs reasoning ? arxiv.org/pdf/2410.05229 [bsky, 0 points, 0 comments]
- Apparently this paper proves and explains that OpenAI can't do 10th grade math Just to put that in context, Deep Blue beat Kasparov for the first time in 1996 Babbage's Difference Engine No. 1 is date [bsky, 0 points, 0 comments]
- New research reveals #LLMs struggle with genuine mathematical #reasoning. Performance varies greatly with small changes, showing sensitivity to irrelevant info and difficulty scaling, and evidencing L [bsky, 0 points, 1 comments]
- @davidaugust Direct link to the paper https://arxiv.org/pdf/2410.05229 (note, not peer-reviewed, yet ?) [bsky, 0 points, 0 comments]
- A paper investigating if LLMs are capable of genuine logical reasoning or just blindly replicating the reasoning found in their data. [bsky, 0 points, 0 comments]
- Understanding the Limitations of Mathematical Reasoning in #LLM "Ultimately, our work underscores significant limitations in the ability of LLMs to perform genuine mathematical reasoning." #ai #rese [bsky, 0 points, 0 comments]
- Ooooh look at me I'm as smart as an Apple Scientist arxiv.org/pdf/2410.05229 TL;DR apple research found that almost all AI benchmarks do well if you ask a question already in their training set, but [bsky, 0 points, 0 comments]
Related