Superhuman performance of a large language model on the reasoning tasks of a physician
2024/12/14 by Peter G. Brodeur, Brodeur, Peter G., Thomas A. Buckley +49 · 36 voices · 13 citations
Medicine · #Artificial Intelligence in Healthcare and Education #Clinical Reasoning and Diagnostic Skills #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2412.10849
openalex publication_date 2024/12/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. Here, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases against a baseline of hundreds of physicians. We conduct five experiments to measure clinical reasoning across differential diagnosis generation, display of diagnostic reasoning, triage differential diagnosis, probabilistic reasoning, and management reasoning, all adjudicated by physician experts with validated psychometrics. We then report a real-world study comparing human expert and AI second opinions in randomly-selected patients in the emergency room of a major tertiary academic medical center in Boston, MA. We compared LLMs and board-certified physicians at three predefined diagnostic touchpoints: triage in the emergency room, initial evaluation by a physician, and admission to the hospital or intensive care unit. In all experiments--both vignettes and emergency room second opinions--the LLM displayed superhuman diagnostic and reasoning abilities, as well as continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have achieved superhuman performance on general medical diagnostic and management reasoning, fulfilling the vision put forth by Ledley and Lusted, and motivating the urgent need for prospective trials.
Cited by
Discussions
- ‼️"o1-preview demonstrates superhuman performance in differential diagnosis, diagnostic clinical reasoning, and management reasoning, superior in multiple domains compared to prior model generations a [bsky, 76 points, 5 comments]
- Superhuman performance of an LLM on the reasoning tasks of a physician [hn, 36 points, 31 comments]
- OpenAI paper comparing LLM to clinician performance on hard tasks. o1 gets superhuman performance as rated by expert (human) clinicians arxiv.org/pdf/2412.10849 [bsky, 12 points, 1 comments]
- Pre-print research finds superhuman performance from LLMs on general medical diagnostic and management reasoning. This is one emergency room, but wow. And that’s just o1-preview. arxiv.org/pdf/2412.10 [bsky, 6 points, 1 comments]
- Recent results on performance of 01-preview model on medical diagnosis & management challenges arxiv.org/abs/2412.10849 For additional reflections, see article on LinkedIn: www.linkedin.com/pulse/adva [bsky, 5 points, 0 comments]
- AI outperformed doctors in medical tasks, identifying the correct diagnosis in 78.3% of cases vs doctors’ 60-70% accuracy. It scored 86% in treatment planning, far exceeding doctors (34%), and achieve [bsky, 5 points, 1 comments]
- This looks like BS. Is it BS? "Superhuman performance of a large language model on the reasoning tasks of a physician" arxiv.org/abs/2412.10849 [bsky, 5 points, 0 comments]
- Superhuman performance of an LLM on the reasoning tasks of a physician [hn, 4 points, 0 comments]
- OpenAI's o1-preview achieves superhuman performance in clinical reasoning: Diagnoses correct in 78.3% of cases, outperforming GPT-4 (88.6% vs 72.9%). Excels in diagnostic & management tasks, scoring [bsky, 4 points, 0 comments]
- One of many things I have been wrong about is that, a few years ago, I would have scoffed at the idea of 'robot doctors'. "We evaluated the medical reasoning abilities of the o1-preview model [findin [bsky, 4 points, 1 comments]
- Can AI match doctors in clinical reasoning? OpenAI's new "o1-preview" model excels in diagnoses & reasoning but struggles with probabilistic tasks. Promising progress, but real-world trials are key to [bsky, 3 points, 0 comments]
- 🩺🤖 New study shows OpenAI's o1 model meets or surpasses human physicians in tackling the toughest diagnostic cases in medicine 🧠📚. It's clear that very soon, *not* using AI could be seen as offeri [bsky, 3 points, 0 comments]
- Superhuman performance of an LLM on the reasoning tasks of a physician [hn, 2 points, 0 comments]
- In a study that included o1-preview (currently 15th on Chatbot Arena*), authors (including @liamgmccoy.bsky.social & @adamrodmanmd.bsky.social) found “In all experiments—both vignettes and emergency [bsky, 2 points, 1 comments]
- One of the newest AI model diagnosed better than "hundreds of expert physicians." This study was out of Harvard: arxiv.org/abs/2412.10849 [bsky, 2 points, 1 comments]
- o1 preview's reported advantage in "diagnostic subtasks" is v. high that clinical decision making supported by models is now a problem worth serious attention. next step would be to document if the di [bsky, 1 points, 0 comments]
- Nothings more fun on a Friday afternoon than reading popular AI paper - Superhuman performance of a large language model on the reasoning tasks of a physician: arxiv.org/pdf/2412.10849 Started with th [bsky, 1 points, 1 comments]
- <3 your work on diagnostic errors. This study shows however the general difficulty in assessing CDSS. 1) It's all about CDSS implementation 2) It's all about the right CDSS (arxiv.org/abs/2412.10849) [bsky, 1 points, 1 comments]
- Pperformance of a large language model on the reasoning tasks of a physician [hn, 1 points, 0 comments]
- Preprint out today that tests o1-preview's medical reasoning experiments against a baseline of 100s of clinicians. In this case the title says it all: Superhuman performance of a large language mod [bsky, 1 points, 0 comments]
- "With continued development, AI may soon become an indispensable assistant to physicians, tackling complex cases with superhuman precision while doctors retain oversight and judgment." [bsky, 1 points, 0 comments]
- Medical diagnostics - quickly ingesting a patient's symptoms and cross-checking it across a massive database of prior patitent conditions - is one areas where LLMs do appear to bring truly game-changi [bsky, 1 points, 1 comments]
- Superhuman performance of a large language model on the reasoning tasks of a physician #llms #diagnostic #physician #clinical #medicine #health #healthcare #acs #o1 [bsky, 1 points, 0 comments]
- https://arxiv.org/abs/2412.10849 大規模言語モデルが医師の推論タスクにおいて、人間を超える性能を発揮したという論文。 LLMの診断能力を医師数百人と比較し、5つの実験と実際の救急現場での比較を行いました。 その結果、LLMはすべての評価項目で医師を上回る「超人的」な診断・推論能力を示し、今後の臨床試験の必要性が示されました。 [bsky, 0 points, 0 comments]
- Superhuman performance of an LLM on the reasoning tasks of a physician https:// arxiv.org/abs/2412.10849 # arxiv # llm [mastodon, 0 points, 0 comments]
- Il suffisait de descendre de 4 réponses pour avoir le lien vers l'étude en question : arxiv.org/abs/2412.10849 [bsky, 0 points, 1 comments]
- Superhuman performance of an LLM on the reasoning tasks of a physician https://arxiv.org/abs/2412.10849 (https://news.ycombinator.com/item?id=44130226) [bsky, 0 points, 0 comments]
- 人工知能情報 医師の推論タスクにおける大規模言語モデルの超人的なパフォーマンス arxiv.org/abs/2412.10849 ChatGTP4o,素人質問 chatgpt.com/share/6777e2... リンクボタンに不具合があっても、url をコピー&ペーストをして、ブラウザで読める可能性があります。また、ChatGTPアプリで開くと読める可能性があります。 実装されてい [bsky, 0 points, 0 comments]
- did you miss it? OpenAI is claiming that o1-preview can outperform experienced physicians (and other LLMs) on complex diagnostic and management tasks, demonstrating a new level of (what they call) “s [bsky, 0 points, 0 comments]
- Superhuman performance of an LLM on the reasoning tasks of a physician https://arxiv.org/abs/2412.10849 [bsky, 0 points, 0 comments]
- Superhuman performance of an LLM on the reasoning tasks of a physician AI model demonstrates advanced medical reasoning capabilities, potentially revolutionizing clinical decision support and diagnos [bsky, 0 points, 0 comments]
- Superhuman performance of a large language model on the reasoning tasks of a physician. arxiv.org/abs/2412.10849 [bsky, 0 points, 0 comments]
- Ref: arxiv.org/pdf/2412.10849 [bsky, 0 points, 0 comments]
- [some-subscribed-rss] New Post: https://arxiv.org/pdf/2412.10849, by Grady https://gwizproductions.info/research/?p=127825 [bsky, 0 points, 0 comments]
- In beiden Fällen schneidet das neue OpenAI LLM o1-preview deutlich besser als andere und eben sogar als menschliche Ärzte mit kliniküblichen Hilfsmitteln (Internet, Literatur, ...) ab. 3/3 P.S. Das w [bsky, 0 points, 0 comments]
- These studies are interesting and keep confirming a few things: 1) LLMs are not better than experts at reasoning 2) Clinical experts using LLMs tend to outperform experts who don't 3) Algorithms push [bsky, 0 points, 0 comments]
Related