The Illusion of Readiness in Health AI
2025/09/22 by Yu‐Cheng Gu, Yu Gu, Jingjing Fu +64 · 13 voices · 4 citations
Decision Sciences · Computer Science · #Scientific Computing and Data Management #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2509.18234
Abstract
Large language models have demonstrated remarkable performance in a wide range of medical benchmarks. Yet underneath the seemingly promising results lie salient growth areas, especially in cutting-edge frontiers such as multimodal reasoning. In this paper, we introduce a series of adversarial stress tests to systematically assess the robustness of flagship models and medical benchmarks. Our study reveals prevalent brittleness in the presence of simple adversarial transformations: leading systems can guess the right answer even with key inputs removed, yet may get confused by the slightest prompt alterations, while fabricating convincing yet flawed reasoning traces. Using clinician-guided rubrics, we demonstrate that popular medical benchmarks vary widely in what they truly measure. Our study reveals significant competency gaps of frontier AI in attaining real-world readiness for health applications. If we want AI to earn trust in healthcare, we must demand more than leaderboard wins and must hold AI systems accountable to ensure robustness, sound reasoning, and alignment with real medical demands.
Citations
Cited by
Discussions
- 👀 www.arxiv.org/abs/2509.18234 [bsky, 45 points, 2 comments]
- Paper: arxiv.org/pdf/2509.18234 [bsky, 13 points, 0 comments]
- The Illusion of Readiness: Stress Testing Frontier Models on Medical Benchmarks [hn, 6 points, 0 comments]
- but I think this critical distance comes in handy when GPT-5 describes in detail an image not provided, complete with chain-of-reasoning “based” on the unprovided photo [bsky, 4 points, 0 comments]
- The Illusion of Readiness: Stress Testing Large Frontier Models on Multimodal Medical Benchmarks www.arxiv.org/abs/2509.18234 🧬🖥️🧪 [bsky, 2 points, 0 comments]
- "If we want AI to earn trust in healthcare, we must demand more than leaderboard wins and must hold AI systems accountable to ensure robustness, sound reasoning, and alignment with real medical demand [bsky, 2 points, 0 comments]
- The illusion of readiness: Stress testing large frontier models on multimodal medical benchmarks www.arxiv.org/abs/2509.18234 #AI [bsky, 1 points, 0 comments]
- No mention of construct validity and no citation to relevant papers (a few below), all while claiming "reasoning". This feels like an indictment of this peer review process. arxiv.org/abs/2503.10694 a [bsky, 1 points, 1 comments]
- The illusion of readiness: Stress testing large frontier models on multimodal medical benchmarks https://www.arxiv.org/abs/2509.18234 #AI [bsky, 0 points, 0 comments]
- Og samtidig har et nyt studie fra Microsoft lige påvist, ingen af de nuværende modeller kan fungere i en medicinsk kontekst. arxiv.org/abs/2509.18234 Ironisk med tanke på, hvor meget Microsoft hypede [bsky, 0 points, 0 comments]
- 医療AIはどうなんだろうね Microsoftからの新しい論文 (ベンチマークが好成績でも、現実世界で医療診断をするには重大な欠陥があるって感じの内容) arxiv.org/abs/2509.182... アンクルジェフの放射線科医は消滅するという10年前の予想に反して、現在なくなってないことについて書かれた投稿なんかも最近バズってたな。いい記事だった。 www.worksinprogress.n [bsky, 0 points, 0 comments]
- arxiv.org/abs/2509.18234 [bsky, 0 points, 1 comments]
- Anyone wants to go to GenAI doctor?? Read this first: www.arxiv.org/abs/2509.18234 [bsky, 0 points, 0 comments]
Related