Comment on The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
2025/06/10 by A. Lawsen, Lawsen, A. · 17 voices · 10 citations
Computer Science · #AI-based Problem Solving and Planning #Constraint Satisfaction and Optimization #Bayesian Modeling and Causal Inference
paper · pdf · doi:10.48550/arxiv.2506.09250
Abstract
Shojaee et al. (2025) report that Large Reasoning Models (LRMs) exhibit "accuracy collapse" on planning puzzles beyond certain complexity thresholds. We demonstrate that their findings primarily reflect experimental design limitations rather than fundamental reasoning failures. Our analysis reveals three critical issues: (1) Tower of Hanoi experiments risk exceeding model output token limits, with models explicitly acknowledging these constraints in their outputs; (2) The authors' automated evaluation framework fails to distinguish between reasoning failures and practical constraints, leading to misclassification of model capabilities; (3) Most concerningly, their River Crossing benchmarks include mathematically impossible instances for N > 5 due to insufficient boat capacity, yet models are scored as failures for not solving these unsolvable problems. When we control for these experimental artifacts, by requesting generating functions instead of exhaustive move lists, preliminary experiments across multiple models indicate high accuracy on Tower of Hanoi instances previously reported as complete failures. These findings highlight the importance of careful experimental design when evaluating AI reasoning capabilities.
Cited by
Discussions
- The Illusion of the Illusion of Thinking – A Comment on Shojaee et al. (2025) [hn, 16 points, 14 comments]
- The Illusion of the Illusion of Thinking [hn, 12 points, 1 comments]
- Comment on the Illusion of Thinking [hn, 4 points, 1 comments]
- arxiv.org/abs/2506.09250 이 논문의 1저자 C. Opus가 언어모델 Claude Opus라고 하는데, Claude라는 이름만 알고 있어서 C. Opus라고 쓰니 뭔지 몰랐음. 따라가기 힘든 세상이다 ... 근데 저걸 저렇게 사람처럼 쓰는 게 이제 괜찮은 건가? [bsky, 3 points, 1 comments]
- Here’s one: “The authors’ evaluation format requires outputting the full sequence of moves at each step, leading to quadratic token growth.” cc: @stellaathena.bsky.social @garymarcus.bsky.social [bsky, 2 points, 1 comments]
- The Illusion of the Illusion of Thinking "The question isn’t whether LRMs can reason, but whether our evaluations can distinguish reasoning from typing." arxiv.org/abs/2506.09250 [bsky, 2 points, 0 comments]
- meanwhile some fucking loser to tried to write a debunk with claude and failed miserably lmao [bsky, 2 points, 0 comments]
- Ha llovido mucho desde que salió este paper. Luego vino The illusion of the illusion of thinking y mil críticas más... arxiv.org/abs/2506.092... [bsky, 2 points, 1 comments]
- Apple's critique of Large Reasoning Models (LRMs) suggests they have limited reasoning abilities, but Anthropic argues this is due to flaws in Apple's testing methods. Anthropic claims that issues lik [bsky, 1 points, 0 comments]
- [1]: arxiv.org/abs/2506.092... [2]: machinelearning.apple.com/research/ill... [3]: x.com/scaling01/st... [bsky, 1 points, 1 comments]
- 「LLMは推論ができない」というのを示すのは難しくて、例えばAppleの論文に対するコメントとして、「そもそも問題が数学的に解けないケースが混ざっていた」「出力トークン数の限界で解けないケースが混ざっていた」といった指摘がある。 arxiv.org/abs/2506.092... また、前記コメントでも触れられているけど、LLMは出力トークン数限界に余裕がある段階でも話をまとめてしまう場合がある。 [bsky, 1 points, 1 comments]
- What are you measuring? How big is a pile of sand? arxiv.org/abs/2506.09250 [bsky, 1 points, 0 comments]
- Kudos for sticking to the bit and having Claude co-author the paper. I'm an AI cynic. I think the recent surge of "reasoning" models is overblown. But yes, Apple's paper is pretty useless. "AI doesn't [bsky, 0 points, 1 comments]
- arxiv.org/pdf/2506.09250 Appleが出した,大規模推論モデルの推論能力にはスケーリング的に限界があるという論文に対する反論が出ているみたいだけれど,ハノイの塔を解くのに何故そんなに苦労するのかよく分からない.私のcopilotのエンドユーザとしての体験だから一般性はないけれど,copilotに何か推論することを要求すると,こちらの感覚としては何かそれらしい言葉は流暢に並べ [bsky, 0 points, 1 comments]
- Comment on The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity https://lobste.rs/s/zeoymy #ai [bsky, 0 points, 0 comments]
- 3/ Enter Alex Lawsen from Open Philanthropy. He just published a brutal counter-study co-authored with Claude Opus itself. Title: "The Illusion of the Illusion of Thinking" His findings tear Apple's r [bsky, 0 points, 0 comments]
- AppleのLLMのreasoning/thinkingはかなり限界があるぞとDisった論文に反論する論文が出ました 筆頭著者 C.Opus (Anthropic) :meow_lol: :meow_lol: :meow_lol: 反論をLLM「が」書いてる https://arxiv.org/abs/2506.09250 [bsky, 0 points, 1 comments]
Related