MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
2024/12/26 by Asma Ben Abacha, Abacha, Asma Ben, Wen-wai Yim +12 · 8 voices · 43 citations
Health Professions · #Algorithm #Artificial intelligence #Benchmark (surveying) #Cartography #Computer science #Electronic Health Records Systems #Error detection and correction #Geography
paper · pdf · doi:10.48550/arxiv.2412.19260
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/12/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Several studies showed that Large Language Models (LLMs) can answer medical questions correctly, even outperforming the average human score in some medical exams. However, to our knowledge, no study has been conducted to assess the ability of language models to validate existing or generated medical text for correctness and consistency. In this paper, we introduce MEDEC (https://github.com/abachaa/MEDEC), the first publicly available benchmark for medical error detection and correction in clinical notes, covering five types of errors (Diagnosis, Management, Treatment, Pharmacotherapy, and Causal Organism). MEDEC consists of 3,848 clinical texts, including 488 clinical notes from three US hospital systems that were not previously seen by any LLM. The dataset has been used for the MEDIQA-CORR shared task to evaluate seventeen participating systems [Ben Abacha et al., 2024]. In this paper, we describe the data creation methods and we evaluate recent LLMs (e.g., o1-preview, GPT-4, Claude 3.5 Sonnet, and Gemini 2.0 Flash) for the tasks of detecting and correcting medical errors requiring both medical knowledge and reasoning capabilities. We also conducted a comparative study where two medical doctors performed the same task on the MEDEC test set. The results showed that MEDEC is a sufficiently challenging benchmark to assess the ability of models to validate existing or generated notes and to correct medical errors. We also found that although recent LLMs have a good performance in error detection and correction, they are still outperformed by medical doctors in these tasks. We discuss the potential factors behind this gap, the insights from our experiments, the limitations of current evaluation metrics, and share potential pointers for future research.
Cited by
- MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
- Search-Based Multi-Trajectory Refinement for Safe C-to-Rust Translation with Large Language Models
- Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records
- A DeepSeek-Powered AI System for Automated Chest Radiograph Interpretation in Clinical Practice
- SocialNav-MoE: A Mixture-of-Experts Vision Language Model for Socially Compliant Navigation with Reinforcement Fine-Tuning
- Ontology Learning with LLMs: A Benchmark Study on Axiom Identification
- Structured Prompting Enables More Robust Evaluation of Language Models
- A Systematic Analysis of Large Language Models with RAG-enabled Dynamic Prompting for Medical Error Detection and Correction
- Jailbreaking Large Vision Language Models in Intelligent Transportation Systems
- Think Before You Retrieve: Learning Test-Time Adaptive Search with Small Language Models
- MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
- Agentic LLMs for REST API Test Amplification: A Comparative Study Across Cloud Applications
- Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World
- ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows
- VAPU: System for Autonomous Legacy Code Modernization
- E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task
- OpenTSLM: Time-Series Language Models for Reasoning over Multivariate Medical Text- and Time-Series Data
- DRES: Benchmarking LLMs for Disfluency Removal
- Large Language Models for Pedestrian Safety: An Application to Predicting Driver Yielding Behavior at Unsignalized Intersections
- A New Benchmark for Evaluating Code Translation with Third-Party Libraries
- Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning
- Empowering Tabular Data Preparation with Language Models: Why and How?
- TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
- Auto-Formulating Dynamic Programming Problems with Large Language Models
- Can Large Language Models Understand As Well As Apply Patent Regulations to Pass a Hands-On Patent Attorney Test?
- Empowering Healthcare Practitioners with Language Models: Structuring Speech Transcripts in Two Real-World Clinical Applications
- RetrySQL: text-to-SQL training with retry data for self-correcting query generation
- MedVAL: Toward Expert-Level Medical Text Validation with Language Models
- Information Loss in LLMs' Multilingual Translation: The Role of Training Data, Language Proximity, and Language Family
- MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports
- Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge
- Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
- TeXpert: A Multi-Level Benchmark for Evaluating LaTeX Code Generation by LLMs
- Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
- Revisiting Uncertainty Estimation and Calibration of Large Language Models
- RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
- OSoRA: Output-Dimension and Singular-Value Initialized Low-Rank Adaptation
- Towards Universal Semantics With Large Language Models
- A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
- Scaling Laws for Moral Machine Judgment in Large Language Models
- FashionM3: Multimodal, Multitask, and Multiround Fashion Assistant based on Unified Vision-Language Model
- LLM-Enhanced Black-Litterman Portfolio Optimization
- A Scoping Review of Natural Language Processing in Addressing Medically Inaccurate Information: Errors, Misinformation, and Hallucination
- SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement Learning
Discussions
- A medical paper from Microsoft lists the previously unknown model sizes of popular closed LLMs - Sonnet3.5: ~175B - GPT3.5-turbo: 175B - GPT4: 1.76T - GPT4o: 200B - GPT4o-mini: 8B - o1-mini: 100B - o1 [bsky, 43 points, 6 comments]
- Microsoft’s latest paper talks about classifying models based on different parameter sizes and shares some rough estimates of the parameter sizes for closed-source models currently in the industry. P [bsky, 10 points, 0 comments]
- Medec: A Benchmark for Medical Error Detection and Correction in Clinical Notes [hn, 2 points, 0 comments]
- This article brings to attention to the less glamorous aspects of LLMs: their domain-specific failures, especially in critical fields such as medicine. While the findings are open to debate, the effor [bsky, 0 points, 0 comments]
- MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes arxiv.org/abs/2412.19260 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2412.19260 [bsky, 0 points, 0 comments]
- MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes A benchmark paper on medical error detection by researchers from Microsoft and UW has attracted attention for including [bsky, 0 points, 0 comments]
- MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes | arXiv arxiv.org/abs/2412.19260 #openaccess [bsky, 0 points, 0 comments]
Related