Can large language models reason about medical questions?
2022/07/17 by Valentin Liévin, Christoffer Egeberg Hother, Christoffer Hother +6 · 2 voices · 86 citations
Computer Science · Medicine · Psychology · #Annotation #Artificial Intelligence in Healthcare and Education #Artificial intelligence #Closing (real estate) #Cognitive psychology #Computer science #Expert system #Machine Learning in Healthcare #Natural language processing #Programming language #Psychology #Recall #Set (abstract data type) #Shot (pellet) #Subject-matter expert #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2207.08143
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2022/07/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Although large language models (LLMs) often produce impressive outputs, it remains unclear how they perform in real-world scenarios requiring strong reasoning skills and expert domain knowledge. We set out to investigate whether close- and open-source models (GPT-3.5, LLama-2, etc.) can be applied to answer and reason about difficult real-world-based questions. We focus on three popular medical benchmarks (MedQA-USMLE, MedMCQA, and PubMedQA) and multiple prompting scenarios: Chain-of-Thought (CoT, think step-by-step), few-shot and retrieval augmentation. Based on an expert annotation of the generated CoTs, we found that InstructGPT can often read, reason and recall expert knowledge. Last, by leveraging advances in prompt engineering (few-shot and ensemble methods), we demonstrated that GPT-3.5 not only yields calibrated predictive distributions, but also reaches the passing score on three datasets: MedQA-USMLE 60.2%, MedMCQA 62.7% and PubMedQA 78.2%. Open-source models are closing the gap: Llama-2 70B also passed the MedQA-USMLE with 62.5% accuracy.
Cited by
- Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
- Evaluating large language models for diagnostic reasoning from unstructured clinical narratives in epilepsy
- Enhancing Clinical Note Generation with ICD-10, Clinical Ontology Knowledge Graphs, and Chain-of-Thought Prompting Using GPT-4
- MedBayes-Lite: A Clinical Uncertainty Governance Layer for Risk-Aware Medical Decision Support
- Multidimensional Rubric-oriented Reward Model Learning via Geometric Projection Reference Constraints
- Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes
- CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
- Large language models encode clinical knowledge
- Fine-Tuned Thoughts: Leveraging Chain-of-Thought Reasoning for Industrial Asset Health Monitoring
- Investigating LLM Capabilities on Long Context Comprehension for Medical Question Answering
- Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents
- Hearing Health in Home Healthcare: Leveraging LLMs for Illness Scoring and ALMs for Vocal Biomarker Extraction
- CURE: Confidence-driven Unified Reasoning Ensemble Framework for Medical Question Answering
- A Survey on Evaluation of Large Language Models
- Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
- MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language Models
- Why Chain of Thought Fails in Clinical Text Understanding
- AgentClinic: a multimodal benchmark for tool-using clinical AI agents
- InterFeat: a pipeline for finding interesting scientific features
- RephQA: Evaluating Readability of Large Language Models in Public Health Question Answering
- MindBenchAI: An Actionable Platform to Evaluate the Profile and Performance of Large Language Models in a Mental Healthcare Context
- A Foundation Model for Chest X-ray Interpretation with Grounded Reasoning via Online Reinforcement Learning
- Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
- Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models
- A Multi-Agent Approach to Neurological Clinical Reasoning
- From EMR Data to Clinical Insight: An LLM-Driven Framework for Automated Pre-Consultation Questionnaire Generation
- Rethinking Evidence Hierarchies in Medical Language Benchmarks: A Critical Evaluation of HealthBench
- Large language models provide unsafe answers to patient-posed medical questions
- Agentic AI framework for End-to-End Medical Data Inference
- ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
- Automating Expert-Level Medical Reasoning Evaluation of Large Language Models
- Multi-Agent Reasoning for Cardiovascular Imaging Phenotype Analysis
- MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
- MIRIAD: Augmenting LLMs with millions of medical query-response pairs
- MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science
- Can Large Language Models Match the Conclusions of Systematic Reviews?
- Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks
- Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Models
- BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science
- Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications
- Aligning Large Language Models with Healthcare Stakeholders: A Pathway to Trustworthy AI Integration
- Generalistic or Specific Embeddings, Which is Better? An Empirical Study on Search for Clinical Coding in Non-English Languages
- MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional
- BRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text
- CheXthought: A global multimodal dataset of clinical chain-of-thought reasoning and visual attention for chest X-ray interpretation
- A Scoping Review of Natural Language Processing in Addressing Medically Inaccurate Information: Errors, Misinformation, and Hallucination
- ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model
Discussions
Related