Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models
2024/06/04 by Marianna Nezhurina, Lucia Cipolina-Kun, Nezhurina, Marianna +5 · 24 voices · 14 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2406.02061
openalex publication_date 2024/06/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large Language Models (LLMs) are often described as being instances of foundation models - that is, models that transfer strongly across various tasks and conditions in few-show or zero-shot manner, while exhibiting scaling laws that predict function improvement when increasing the pre-training scale. These claims of excelling in different functions and tasks rely on measurements taken across various sets of standardized benchmarks showing high scores for such models. We demonstrate here a dramatic breakdown of function and reasoning capabilities of state-of-the-art models trained at the largest available scales which claim strong function, using a simple, short, conventional common sense problem (AIW problem) formulated in concise natural language, easily solvable by humans. The breakdown is dramatic, as models show strong fluctuations across even slight problem variations that should not affect problem solving, also expressing strong overconfidence in the wrong solutions, often backed up by plausible sounding explanation-like confabulations. Various standard interventions in an attempt to get the right solution, like various type of enhanced prompting, or urging the models to reconsider the wrong solutions again by multi step re-evaluation, fail. We take these initial observations to the scientific and technological community to stimulate urgent re-assessment of the claimed capabilities of current generation of LLMs. Such re-assessment also requires common action to create standardized benchmarks that would allow proper detection of such basic reasoning deficits that obviously manage to remain undiscovered by current state-of-the-art evaluation procedures and benchmarks. Code for reproducing experiments in the paper and raw experiments data can be found at https://github.com/LAION-AI/AIW
Cited by
Discussions
- Simple tasks showing reasoning breakdown in state-of-the-art LLMs [hn, 375 points, 380 comments]
- Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models [lobsters, 26 points, 14 comments]
- [Paper] Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in SOTA Large Language Models [lemmy, 24 points, 4 comments]
- Like the AIW question suite for unimodal LLMs are in some sense “puzzles”. VLMs fail at much lower complexity than that [bsky, 3 points, 0 comments]
- Nice paper on LLMs struggle to solve simple reasoning tasks. I love comment "What becomes quite clear through our study is the failure of current standardized benchmarks to reflect true model reasoni [bsky, 3 points, 0 comments]
- I don't know how to break this to commentators covering machine learning topics, but LLMs do not "reason", ever, so they cannot suffer from a breakdown in reasoning. arxiv.org/abs/2406.02061 [bsky, 2 points, 0 comments]
- IMNSHO LLMs are fundamentally ill-conceived. Without the infrastructure of more primitive affairs like signs, meaning, reference etc. their playing with word statistics is useless. Hence arxiv.org/abs [bsky, 1 points, 0 comments]
- 🔎 Alice in Wonderland gives state-of-the-art large language models a 'complete reasoning breakdown'. #technology arxiv.org/abs/2406.02061 [bsky, 1 points, 0 comments]
- Here’s a longer paper but with relatively easy to understand examples: arxiv.org/pdf/2406.02061 Here’s another paper by Apple along these lines: arxiv.org/pdf/2410.05229 Here’s an article from one o [bsky, 1 points, 0 comments]
- And here's a link to the paper. arxiv.org/abs/2406.02061 [bsky, 1 points, 1 comments]
- Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in LLMs (arxiv.org) Main Link | Discussion [bsky, 1 points, 0 comments]
- arxiv.org/abs/2406.02061 [bsky, 1 points, 0 comments]
- This is a fascinating study showing how badly the Large Language Models (LLMs) like GPT are at solving simple problems. arxiv.org/abs/2406.02061 [bsky, 1 points, 0 comments]
- You think AI replaces cognition? It can’t even solve the most rudimentary problems: arxiv.org/abs/2406.02061 [bsky, 1 points, 0 comments]
- AGI is a bubble waiting to pop. "Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models" arxiv.org/abs/2406.02061 [bsky, 1 points, 0 comments]
- Simple tasks showing reasoning breakdown in state-of-the-art LLMs https://arxiv.org/abs/2406.02061 [bsky, 0 points, 0 comments]
- This looks interesting. arxiv.org/abs/2406.02061 [bsky, 0 points, 0 comments]
- 最近の性能の高いLLMでも“Alice has N brothers and she also has M sisters. How many sisters does Alice’s brother have?“という問題に正しく答えられないという論文。 arxiv.org/abs/2406.02061 [bsky, 0 points, 1 comments]
- " models also express strong overconfidence in their wrong solutions, while providing often non-sensical "reasoning"-like explanations akin to confabulations to justify and backup the validity of thei [bsky, 0 points, 0 comments]
- arxiv.org/abs/2406.02061. - LLMs are bad at things like reasoning? Really? Who could've guessed?!? [bsky, 0 points, 0 comments]
- What happens when you ask a LLM the question, "Alice has N brothers and she also has M sisters. How many sisters does Alice’s brother have?" where N is, say, 4 and M is 1? It frequently breaks, and b [bsky, 0 points, 0 comments]
- But as soon as you get to the edge of space or outside it (rare combinations of words in the training set), everything start to fall apart. Also it's interesting how bad LLMs are once you ask them to [bsky, 0 points, 0 comments]
- There are real, reliably reproducible problems that pass the checks. For example: https://arxiv.org/abs/2406.02061 But they're a lot harder to find than you'd think, looking at social media. [bsky, 0 points, 0 comments]
- Solange die LLMs durch die Bank an „Alice hat zwei Schwestern und vier Brüder. Wieviele Schwestern hat Alices Bruder?“ scheitern, würde ich darüber noch keinen Schlaf verlieren. Die Dinger haben halt [bsky, 0 points, 1 comments]
Related