Sequential Diagnosis with Language Models
2025/06/27 by Harsha Nori, Mayank Daswani, Nori, Harsha +27 · 15 voices · 18 citations
#cs.CL
paper · pdf · doi:10.48550/arxiv.2506.22405
Abstract
Artificial intelligence holds great promise for expanding access to expert medical knowledge and reasoning. However, most evaluations of language models rely on static vignettes and multiple-choice questions that fail to reflect the complexity and nuance of evidence-based medicine in real-world settings. In clinical practice, physicians iteratively formulate and revise diagnostic hypotheses, adapting each subsequent question and test to what they've just learned, and weigh the evolving evidence before committing to a final diagnosis. To emulate this iterative process, we introduce the Sequential Diagnosis Benchmark, which transforms 304 diagnostically challenging New England Journal of Medicine clinicopathological conference (NEJM-CPC) cases into stepwise diagnostic encounters. A physician or AI begins with a short case abstract and must iteratively request additional details from a gatekeeper model that reveals findings only when explicitly queried. Performance is assessed not just by diagnostic accuracy but also by the cost of physician visits and tests performed. We also present the MAI Diagnostic Orchestrator (MAI-DxO), a model-agnostic orchestrator that simulates a panel of physicians, proposes likely differential diagnoses and strategically selects high-value, cost-effective tests. When paired with OpenAI's o3 model, MAI-DxO achieves 80% diagnostic accuracy--four times higher than the 20% average of generalist physicians. MAI-DxO also reduces diagnostic costs by 20% compared to physicians, and 70% compared to off-the-shelf o3. When configured for maximum accuracy, MAI-DxO achieves 85.5% accuracy. These performance gains with MAI-DxO generalize across models from the OpenAI, Gemini, Claude, Grok, DeepSeek, and Llama families. We highlight how AI systems, when guided to think iteratively and act judiciously, can advance diagnostic precision and cost-effectiveness in clinical care.
Citations
Cited by
Discussions
- Agentic A.I. vs experienced physicians for diagnosis of > 300 complex diagnostic cases: 4-fold higher accuracy and 20% lower cost arxiv.org/abs/2506.22405 [bsky, 83 points, 10 comments]
- Microsoft announces LLM ensemble model that correctly diagnoses 85% of cases in the New England Journal of Medicine clinicopathological conference cases. 4x higher than the average of generalist physi [bsky, 9 points, 0 comments]
- Bad evaluation, sounds like the AI tool could search for the answers but the humans could not. From the paper: "possible that some off-the-shelf models were trained on [part of the test set but] ... w [bsky, 7 points, 0 comments]
- While early studies showing #genAI passing MD license/board tests & "showing" empathy were cool, the @Google AMIE study & new @Microsoft diagnosis study arxiv.org/abs/2506.22405 make clear that AI wil [bsky, 7 points, 1 comments]
- Sequential Diagnosis with Language Models [hn, 4 points, 0 comments]
- Here is the paper. I look forward to eviscerating, er, I mean reviewing, it later. arxiv.org/pdf/2506.22405 [bsky, 4 points, 1 comments]
- Thanks for posting this terrific paper from @erichorvitz.bsky.social and team! Neurologists will be on board as we traditionally embrace *sequential diagnosis* where each finding generates a new diffe [bsky, 3 points, 1 comments]
- Microsoft: Sequential Diagnosis with Language Models [hn, 2 points, 0 comments]
- Sequential Diagnosis with Language Models [hn, 2 points, 1 comments]
- Strange that Microsoft AI team don’t seem to understand that NEJM cases are often diverse, difficult and obscure and treatment is typically specialist led, which is NOT what they tested. They handicap [bsky, 2 points, 1 comments]
- Sequential Diagnosis with Language Models #md #physician #diagnostic #cost #medicine #health #healthcare #maidxo #microsoft #llms #artificialintelligence [bsky, 1 points, 0 comments]
- Microsoft researchers transformed 304 complex NEJM cases into interactive diagnosis challenges and found that their AI system, MAI-DxO, achieved 4× higher diagnostic accuracy than physicians while cut [bsky, 1 points, 0 comments]
- "Benchmarked against real-world case records published each week in the New England Journal of Medicine, we show that Microsoft AI Diagnostic Orchestrator (MAI-DxO) correctly diagnoses up to 85% of NE [bsky, 0 points, 1 comments]
- #MedSky #MLSky Direct link to the pre-print: arxiv.org/abs/2506.22405 [bsky, 0 points, 0 comments]
- More accurate than clinicians [bsky, 0 points, 0 comments]
Related