Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
2025/08/02 by Chengshuai Zhao, Zhao, Chengshuai, Zhen Tan +13 · 35 voices · 20 citations
#cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2508.01191
Abstract
Chain-of-Thought (CoT) prompting has been shown to be effective in eliciting structured reasoning (i.e., CoT reasoning) from large language models (LLMs). Regardless of its popularity, recent studies expose its failures in some reasoning tasks, raising fundamental questions about the nature of CoT reasoning. In this work, we propose a data distribution lens to understand when and why CoT reasoning succeeds or fails. We hypothesize that CoT reasoning reflects a structured inductive bias learned from in-distribution data, enabling models to conditionally generate reasoning trajectories that approximate those observed during training. As such, the effectiveness of CoT reasoning is fundamentally governed by the nature and degree of distribution discrepancy between training data and test queries. Guided by this lens, we dissect CoT reasoning via three dimensions: task, length, and format. To test the hypothesis, we introduce DataAlchemy, an abstract and fully controllable environment that trains LLMs from scratch and systematically probes them under various distribution conditions. Through rigorous controlled experiments, we reveal that CoT reasoning is a brittle mirage when it is pushed beyond training distributions, emphasizing the ongoing challenge of achieving genuine and generalizable reasoning.
Citations
Cited by
Discussions
- New paper reveals Chain-of-Thought reasoning of LLMs a mirage [lemmy, 31 points, 0 comments]
- "Our results reveal that [Chain-of-Thought] reasoning is a brittle mirage that vanishes when it is pushed beyond training distributions" [bsky, 4 points, 0 comments]
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens [hn, 4 points, 0 comments]
- 📢 Excited to share our new arXiv preprint: "Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens" 📄 Read the full paper: arxiv.org/abs/2508.01191 💻 Code & data: github.com/Cheng [bsky, 3 points, 15 comments]
- "Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens": arxiv.org/pdf/2508.01191 "Together, these findings suggest that LLMs are not principled reasoners but rather sophisticated s [bsky, 3 points, 1 comments]
- Is Chain-of-Thought Reasoning of LLMs a Mirage? [hn, 3 points, 0 comments]
- arxiv.org/pdf/2508.01191 Here you go! [bsky, 3 points, 0 comments]
- Is Chain-of-Thought Reasoning of LLMs a Mirage? [hn, 3 points, 0 comments]
- Chain of Thought reasoning is “a brittle mirage that vanishes when it is pushed beyond training distributions” arxiv.org/pdf/2508.01191 [bsky, 2 points, 0 comments]
- Viral AI paper of the day: "Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens" arxiv.org/pdf/2508.01191 It's often getting headlines like "LLMs' 'Simulated Reasoning' Abilities [bsky, 2 points, 1 comments]
- Interesting paper on the likely brittleness of "chain-of-thought reasoning" of LLMs. The framework they use to generate data for controlled experiments is cool, too. arxiv.org/abs/2508.01191 [bsky, 1 points, 0 comments]
- Do they think or give a convincing illusion of thinking? arxiv.org/pdf/2508.01191 [bsky, 1 points, 1 comments]
- New paper: "Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens" "Our results reveal that [Chain of Thought] reasoning is a brittle mirage that vanishes when it is pushed beyond t [bsky, 1 points, 0 comments]
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens [hn, 1 points, 0 comments]
- Oh, pointer to a good paper on the topic of the faux reasoning and explanations. When people earnestly say "try harder" I like to give them papers. Betteridge's law of headlines applies to this. arxiv [bsky, 0 points, 0 comments]
- New cool paper making the rounds on Twitter! Will try to track down & tag any authors that exist on Bluesky, if I can find them, after I finish reading it: arxiv.org/pdf/2508.01191 (tag them for me, i [bsky, 0 points, 0 comments]
- @edzitron.com arxiv.org/pdf/2508.01191 [bsky, 0 points, 0 comments]
- Preprint: "Is Chain-of-Thought Reasoning of #LLMs a Mirage? A Data Distribution Lens" (via #arXiv) arxiv.org/abs/2508.01191 #AI [bsky, 0 points, 0 comments]
- arxiv.org/pdf/2508.01191 [bsky, 0 points, 0 comments]
- Sam Altman's AI marketing claims are bullshit. His and other LLM vendor offerings do not *reason* like we sentient animals; they are capable of no more than sophisticated but essentially dumb pattern [bsky, 0 points, 0 comments]
- The leap year example is out of a scientific paper but of course models change one day to the next: www.arxiv.org/abs/2508.01191 [bsky, 0 points, 1 comments]
- "This work offers a deeper understanding of why and when CoT reasoning fails, emphasizing the ongoing challenge of achieving genuine and generalizable reasoning." arxiv.org/pdf/2508.01191 [bsky, 0 points, 0 comments]
- I don't know - are humans really that much better outside of their training data? Isn't this similar to making educated guesses and trial and error in problem-solving? Sometimes we get it wrong. arxiv [bsky, 0 points, 0 comments]
- Betteridge’s Law of headlines states that any headline in the form of a question can be answered with the word “no.” I’m beginning to believe that any academic paper’s title in the form of a question [bsky, 0 points, 0 comments]
- David Gerard linked this paper in one of his articles about LLMs and it's very interesting so far: arxiv.org/abs/2508.01191 essentially, all that "chain of reasoning" stuff the new AI models are pushi [bsky, 0 points, 1 comments]
- https://arxiv.org/pdf/2508.01191 "Together, these findings suggest that LLMs are not principled reasoners but rather sophisticated simulators of reasoning-like text.” Why spend energy examining the so [bsky, 0 points, 0 comments]
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens arxiv.org/pdf/2508.01191 [bsky, 0 points, 0 comments]
- [Chain of thought] reasoning effectively reproduces reasoning patterns closely aligned with training distributions but suffers significant degradation when faced with distributional deviations. Such o [bsky, 0 points, 0 comments]
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens https://arxiv.org/pdf/2508.01191 [bsky, 0 points, 0 comments]
- It's basically the main point of Zhao et al (arxiv.org/pdf/2508.01191), and my point is mostly that it introduces a fundamental unreliability in said outputs, *especially* when they most closely resem [bsky, 0 points, 1 comments]
- “LLMs are not principled reasoners but rather sophisticated simulators of reasoning-like text” arxiv.org/abs/2508.01191 [bsky, 0 points, 0 comments]
- new study exposes the illusion of chain-of-thought reasoning in llms: implications for ai entrepreneurship in europe arxiv.org [bsky, 0 points, 0 comments]
- "Is Chain-of-Thought Reasoning of LLMs a Mirage?": "Chain-of-Thought (CoT) prompting has been shown to improve Large Language Model (LLM) performance on various tasks. With this approach, LLMs appear [bsky, 0 points, 0 comments]
- Is Chain-of-thought and “reasoning” merely a simulation of reasoning, and therefore constrained by pre-training data? arxiv.org/pdf/2508.01191 [bsky, 0 points, 1 comments]
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens #llms #reasoning #generativeai #cot #chainofthought #dataalchemy [bsky, 0 points, 0 comments]
Related