MIRAGE: The Illusion of Visual Understanding
2026/03/23 by Mohammad Asadi, Jack W. O'Sullivan, Fang Cao +5 · 45 voices · 6 citations
#cs.AI
paper · pdf
Abstract
Multimodal AI systems have achieved remarkable performance across a broad range of real-world tasks, yet the mechanisms underlying visual-language reasoning remain surprisingly poorly understood. We report three findings that challenge prevailing assumptions about how these systems process and integrate visual information. First, Frontier models readily generate detailed image descriptions and elaborate reasoning traces, including pathology-biased clinical findings, for images never provided; we term this phenomenon mirage reasoning. Second, without any image input, models also attain strikingly high scores across general and medical multimodal benchmarks, bringing into question their utility and design. In the most extreme case, our model achieved the top rank on a standard chest X-ray question-answering benchmark without access to any images. Third, when models were explicitly instructed to guess answers without image access, rather than being implicitly prompted to assume images were present, performance declined markedly. Explicit guessing appears to engage a more conservative response regime, in contrast to the mirage regime in which models behave as though images have been provided. These findings expose fundamental vulnerabilities in how visual-language models reason and are evaluated, pointing to an urgent need for private benchmarks that eliminate textual cues enabling non-visual inference, particularly in medical contexts where miscalibrated AI carries the greatest consequence. We introduce B-Clean as a principled solution for fair, vision-grounded evaluation of multimodal AI systems.
Citations
Cited by
Discussions
- There’s some cool recent research on this phenomenon! It turns out vision language models excel at image benchmarks *even when the actual images aren’t provided,* because the answers are implicit in t [bsky, 235 points, 10 comments]
- After discovering a recent article about LLMs fabricating analyses from nonexistent data (arxiv.org/abs/2603.21687), I did an experiment with our new lab-sanctioned chatbot. Well, here’s the result (i [bsky, 79 points, 5 comments]
- "In the most extreme case, our model achieved the top rank on a standard chest X-ray question-answering benchmark without access to any images." Well this isn't good 😮 arxiv.org/abs/2603.21687 [bsky, 61 points, 2 comments]
- Stanford study reveals AI vision models invent images they never see [hn, 49 points, 1 comments]
- Oh gosh. Remember the excitement about frontier models understanding medical images and acing tests? "In the most extreme case, our model achieved the top rank on a standard chest X-ray question-answe [bsky, 45 points, 5 comments]
- Preprint, so take it with a grain of salt, but this is a funny one: "In the most extreme case, our model achieved the top rank on a standard chest X-ray question-answering benchmark without access to [bsky, 25 points, 1 comments]
- Preprint, so take it with a grain of salt, but this is a funny one: "In the most extreme case, our model achieved the top rank on a standard chest X-ray question-answering benchmark without access to [bsky, 25 points, 1 comments]
- Moderne multimodale LLMs erzeugen den Eindruck echten Bildverständnisses. Aber sie extrapolieren nur aus trainierten Textmustern und statistischem Wissen. Sie erzeugen plausible Bildbeschreibungen, di [bsky, 14 points, 1 comments]
- Mirage Reasoning: The Illusion of Visual Understanding [hn, 12 points, 4 comments]
- this is an established problem - it's called mirage reasoning - the slop generated is independent of whether it has the imagery or in his case sound files at all arxiv.org/pdf/2603.21687 [bsky, 6 points, 0 comments]
- arxiv.org/abs/2603.21687 spotted in the wild [bsky, 5 points, 0 comments]
- (Long thread, thought dumping) About a few weeks ago I found a paper about AI research, and I wanted to let a few thoughts out because it was shocking to me in a new way, despite already being vocal a [bsky, 5 points, 1 comments]
- Har du også hørt at KI gjør at man setter mye bedre diagnoser? Vel, noen forskere har sett på LLM-er som brukes til å gi diagnose på grunnlag av bilder. Men de glemte å gi LLM-en tilgang til bildene, [bsky, 4 points, 0 comments]
- "Mirage: The Illusion of Visual Understanding" arxiv.org/abs/2603.21687 New research shows that multimodal LLMs may base their answer/reasoning on a non-existent image. Includes a new eval framework B [bsky, 4 points, 0 comments]
- Actually it turns out that while machine vision is good for cancer detection, this "AI" LLM related software isn't useful for than. It looked like it was in a couple studies, but that's just because i [bsky, 4 points, 1 comments]
- 'Despite the responses being identical for the UK and US, Copilot produced a rich, detailed summary of how US and UK respondents differed' Nice post on the danger of careless use of #AI for text and d [bsky, 4 points, 0 comments]
- They "know" so much that they'll give me pages of legal contract review with "direct quotes" from a fake filename of a screenshot of a contract that does not exist. These are construct validity failur [bsky, 4 points, 1 comments]
- Inspired by the Mirage paper¹, I asked ChatGPT to identify an image I did not actually upload. It described this image in some detail but said that the text was to blurry. Luckily, it could upscale t [bsky, 3 points, 1 comments]
- Did you see Euan Ashley’s new paper? arxiv.org/abs/2603.21687 [bsky, 3 points, 2 comments]
- FWIW, visual-language reasoning holed below the waterline in this 23 March 2026 paper. VLR generates detailed image descriptions and elaborate reasoning traces even in the absense of images; VLR bench [bsky, 3 points, 0 comments]
- והיום בפינת clever llm >> clever Hans, הידועה גם כ"קשה לדעת שאתה בודק מה שאתה חושב שאתה בודק, ושבעתיים בllmים": מודלים לניתוח תמונות עברו בנצ'מרק בציונים גבוהים *מבלי שסופקו להם התמונות*, ככהנ בהתבסס [bsky, 2 points, 1 comments]
- [2603.21687] MIRAGE: The Illusion of Visual Understanding — Worth a quick read if you work with vision-language models: MIRAGE argues that apparent “visual understanding” can be an illusion driven by [bsky, 1 points, 0 comments]
- 🟠 Stanford/Fei-Fei Li Paper: LLMs Beat Radiologists on Medical Image Benchmarks — Without Seeing Any Images A Stanford study co-authored by Fei-Fei Li reveals that frontier LLMs achieve top-ranked pe [bsky, 1 points, 0 comments]
- a Stanford/Fei-Fei Li team fed frontier models a chest X-ray benchmark with no images attached. the models still generated detailed pathology findings and scored top rank. they call it mirage reasonin [bsky, 1 points, 0 comments]
- MIRAGE: The Illusion of Visual Understanding arxiv.org/abs/2603.21687 [bsky, 1 points, 0 comments]
- Re benchmarks, reminds me of this article... arxiv.org/abs/2603.21687 [bsky, 1 points, 1 comments]
- "Multimodal #AI systems have achieved remarkable performance across a broad range of real-world tasks, yet the mechanisms underlying visual-language reasoning remain surprisingly poorly understood." # [bsky, 1 points, 0 comments]
- AI evaluates in its stride a medical image it has not been given. arxiv.org/abs/2603.21687 [bsky, 1 points, 0 comments]
- "Mirage: The Illusion of Visual Understanding" https://arxiv.org/abs/2603.21687 New research shows that multimodal LLMs may base their answer/reasoning on a non-existent image. Includes a new eval fra [bsky, 1 points, 1 comments]
- …see this (arxiv.org/pdf/2603.21687) and this (kucharski.substack.com/p/real-signa...), for example. I'm also very much not convinced by the argument that a book wouldn't otherwise be translated into [bsky, 1 points, 1 comments]
- Stanford Chair of Medicine: LLMs Are Superhuman Guessers [lemmy, 1 points, 0 comments]
- MIRAGE: THE ILLUSION OF VISUAL UNDERSTANDING: arxiv.org/pdf/2603.21687 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2603.21687 [bsky, 0 points, 0 comments]
- Building for test driven development is kind of the main software engineering task. As the enemy, the computer doing its best to convince you it’s fine when it’s not is its job. Killing off the “monad [bsky, 0 points, 0 comments]
- saw this paper this morning which I thought was interesting, specifically on the way in which benchmarking might be flawed (this specifically wrt medical usecases) arxiv.org/abs/2603.21687 [bsky, 0 points, 1 comments]
- Holy Crap. #Ai #Fail #vision #mirage [bsky, 0 points, 0 comments]
- arxiv.org/abs/2603.21687 MIRAGE: The Illusion of Visual Understanding [bsky, 0 points, 0 comments]
- MIRAGE: The illusion of visual understanding (by AI models) arxiv.org/abs/2603.21687 #LLM #AI [bsky, 0 points, 0 comments]
- MIRAGE: The illusion of visual understanding (by AI models) https://arxiv.org/abs/2603.21687 #LLM #AI [bsky, 0 points, 0 comments]
- arxiv.org/abs/2603.21687 この論文の音声解説を作ったらすごく面白かった。 [bsky, 0 points, 0 comments]
- MIRAGE: the illusion of visual understanding (@arxiv.org) Main Link | Discussion [bsky, 0 points, 0 comments]
- Stanford study reveals AI vision models invent images they never see https://arxiv.org/abs/2603.21687 (https://news.ycombinator.com/item?id=47570650) [bsky, 0 points, 0 comments]
- RT @euanashley: New AI paper from us this week. When my student first showed me his initial findings, I really didn’t know what to make of them. I felt that this was an interesting but curious loophol [bsky, 0 points, 0 comments]
- "...we show that multimodal AI systems can appear to see when they do not, reason about images that were never provided, and achieve high benchmark scores without genuine visual access." arxiv.org/abs [bsky, 0 points, 0 comments]
- Yikes: "Frontier models readily generate detailed image descriptions and elaborate reasoning traces, including pathology-biased clinical findings, for images never provided; we term this phenomenon mi [bsky, 0 points, 0 comments]
Related