What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
2020/09/28 by Di Jin, Eileen Pan, Jin, Di +10 · 1 voice · 442 citations
Computer Science · #Natural Language Processing Techniques #Sentiment Analysis and Opinion Mining #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2009.13081
Submitted to AAAI 2021
arxiv created 2020/09/28 · arxiv published 2020/09/28 · arxiv updated 2020/09/29
Abstract
Open domain question answering (OpenQA) tasks have been recently attracting more and more attention from the natural language processing (NLP) community. In this work, we present the first free-form multiple-choice OpenQA dataset for solving medical problems, MedQA, collected from the professional medical board exams. It covers three languages: English, simplified Chinese, and traditional Chinese, and contains 12,723, 34,251, and 14,123 questions for the three languages, respectively. We implement both rule-based and popular neural methods by sequentially combining a document retriever and a machine comprehension model. Through experiments, we find that even the current best method can only achieve 36.7%, 42.0%, and 70.1% of test accuracy on the English, traditional Chinese, and simplified Chinese questions, respectively. We expect MedQA to present great challenges to existing OpenQA systems and hope that it can serve as a platform to promote much stronger OpenQA models from the NLP community in the future.
Cited by
- Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference
- MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA
- Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings
- Auditing Institutional Heterogeneity for Generative AI in Patient Education: A Large-Scale Study of 102 US Transplant Handbooks
- What Does Chain-of-Thought Contribute at Probe Time? Evidence for Local Co-Occurrence Activation
- MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models
- Evaluating large language models for diagnostic reasoning from unstructured clinical narratives in epilepsy
- VeriLLMed: Interactive Visual Debugging of Medical Large Language Models with Knowledge Graphs
- MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
- Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
- Automatic Replication of LLM Mistakes in Medical Conversations
- A Real-World Evaluation of LLM Medication Safety Reviews in NHS Primary Care
- Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight
- QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation
- MoE Pathfinder: Trajectory-driven Expert Pruning
- ReX-MLE: The Autonomous Agent Benchmark for Medical Imaging Challenges
- DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
- Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
- Bolmo: Byteifying the Next Generation of Language Models
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- Multi-Agent Medical Decision Consensus Matrix System: An Intelligent Collaborative Framework for Oncology MDT Consultations
- MedCEG: Reinforcing Verifiable Medical Reasoning with Critical Evidence Graph
- Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
- AI Benchmark Democratization and Carpentry
- Script Gap: Evaluating LLM Triage on Indian Languages in Native vs Romanized Scripts in a Real World Setting
- CP-Env: Evaluating Large Language Models on Clinical Pathways in a Controllable Hospital Environment
- AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding
- Cooperative Retrieval-Augmented Generation for Question Answering: Mutual Information Exchange and Ranking by Contrasting Layers
- MedBioRAG: Semantic Search and Retrieval-Augmented Generation with Large Language Models for Medical and Biological QA
- MentraSuite: Post-Training Large Language Models for Mental Health Reasoning and Assessment
- A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
- Metric-Fair Prompting: Treating Similar Samples Similarly
- Multilingual Medical Reasoning for Question Answering with Large Language Models
- MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis
- E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing
- Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks
- Clinical-R1: Empowering Large Language Models for Faithful and Comprehensive Reasoning with Clinical Objective Relative Policy Optimization
- OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
- Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
- Mirror, Mirror on the Wall -- Which is the Best Model of Them All?
- MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology
- InvisibleBench: A Deployment Gate for Caregiving Relationship AI
- A Systematic Analysis of Large Language Models with RAG-enabled Dynamic Prompting for Medical Error Detection and Correction
- Knowing How to Edit: Reliable Evaluation Signals for Diagnosing and Optimizing Prompts at Query Level
- Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
- SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
- APRIL: Annotations for Policy evaluation with Reliable Inference from LLMs
- Fantastic Bugs and Where to Find Them in AI Benchmarks
- HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning
- GPS: General Per-Sample Prompter
- Let the Model Distribute Its Doubt: Confidence Estimation through Verbalized Probability Distribution
- Can Large Language Models Function as Qualified Pediatricians? A Systematic Evaluation in Real-World Clinical Contexts
- A Novel Hierarchical Integration Method for Efficient Model Merging in Medical LLMs
- MedSumGraph: enhancing GraphRAG for medical QA with summarization and optimized prompts
- Assessing Automated Fact-Checking for Medical LLM Responses with Knowledge Graphs
- 47B Mixture-of-Experts Beats 671B Dense Models on Chinese Medical Examinations
- Expert-Guided Prompting and Retrieval-Augmented Generation for Emergency Medical Service Question Answering
- Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
- iMAD: Intelligent Multi-Agent Debate for Efficient and Accurate LLM Inference
- Beyond Elicitation: Provision-based Prompt Optimization for Knowledge-Intensive Tasks
- Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning
- EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
- SlideBot: A Multi-Agent Framework for Generating Informative, Reliable, Multi-Modal Presentations
- Human or LLM as Standardized Patients? A Comparative Study for Medical Education
- Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
- Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
- Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
- What Makes Reasoning Invalid: Echo Reflection Mitigation for Large Language Models
- Overview of CHIP 2025 Shared Task 2: Discharge Medication Recommendation for Metabolic Diseases Based on Chinese Electronic Health Records
- CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
- MIMIC-SR-ICD11: A Dataset for Narrative-Based Diagnosis
- An MLCommons Scientific Benchmarks Ontology
- KGFR: A Foundation Retriever for Generalized Knowledge Graph Question Answering
- RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation
- CareMedEval dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field
- LA-MARRVEL: A Knowledge-Grounded, Language-Aware LLM Framework for Clinically Robust Rare Disease Gene Prioritization
- Demo: Statistically Significant Results On Biases and Errors of LLMs Do Not Guarantee Generalizable Results
- MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
- Evontree: Ontology Rule-Guided Self-Evolution of Large Language Models
- The Geometry of Dialogue: Graphing Language Models to Reveal Synergistic Teams for Multi-Agent Collaboration
- EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis
- Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
- Before Agents Speak: Pre-hoc Failure Risk Inference in Multi-Agent Systems
- SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning
- GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring
- Automated Composition of Agents: A Knapsack Approach for Agentic Component Selection
- Dissecting Role Cognition in Medical LLMs via Neuronal Ablation
- Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives
- Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
- Towards Transparent Reasoning: What Drives Faithfulness in Large Language Models?
- META-RAG: Meta-Analysis-Inspired Evidence-Re-Ranking Method for Retrieval-Augmented Generation in Evidence-Based Medicine
- Towards Scalable Oversight via Partitioned Human Supervision
- Frustratingly Easy Task-aware Pruning for Large Language Models
- Weak-to-Strong Generalization under Distribution Shifts
- Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs
- Investigating LLM Capabilities on Long Context Comprehension for Medical Question Answering
- Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency
- IMB: An Italian Medical Benchmark for Question Answering
- Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language Models
- From Retrieval to Generation: Unifying External and Parametric Knowledge for Medical Question Answering
- From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
- Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- Demo: Guide-RAG: Evidence-Driven Corpus Curation for Retrieval-Augmented Generation in Long COVID
- MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question Answering
- CURE: Confidence-driven Unified Reasoning Ensemble Framework for Medical Question Answering
- GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians
- Big Reasoning with Small Models: Instruction Retrieval at Inference Time
- BioMedSearch: A Multi-Source Biomedical Retrieval Framework Based on LLMs
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- MedKGEval: A Knowledge Graph-Based Multi-Turn Evaluation Framework for Open-Ended Patient Interactions with Clinical LLMs
- Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs
- CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
- AgentCaster: Reasoning-Guided Tornado Forecasting
- Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks
- Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
- Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
- PISA: A Pragmatic Psych-Inspired Unified Memory System for Enhanced AI Agency
- LONGQAEVAL: Designing Reliable Evaluations of Long-Form Clinical QA under Resource Constraints
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models
- ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled Tuning
- Between Knowledge and Care: Evaluating Generative AI-Based IUI in Type 2 Diabetes Management Through Patient and Physician Perspectives
- Learning Reasoning Rewards from Expert Demonstrations with Inverse Reinforcement Learning
- IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
- Don't Throw Away Your Pretrained Model
- DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
- CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts
- Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation
- Understanding the Effects of Domain Finetuning on LLMs
- ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
- Recent Advances in Automated Question Answering In Biomedical Domain
- CaRT: Teaching LLM Agents to Know When They Know Enough
- MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding
- NurseLLM: The First Specialized Language Model for Nursing
- YpathRAG:A Retrieval-Augmented Generation Framework and Benchmark for Pathology
- FedSRD: Sparsify-Reconstruct-Decompose for Communication-Efficient Federated Large Language Models Fine-Tuning
- Evaluation of Clinical Trials Reporting Quality using Large Language Models
- Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement Learning
- MedReflect: Teaching Medical LLMs to Self-Improve via Reflective Correction
- Operationalizing Data Minimization for Privacy-Preserving LLM Prompting
- Triplet-Structured Knowledge Integration for Multi-Turn Medical Reasoning
- Rethinking Reward Models for Multi-Domain Test-Time Scaling
- Automated Evaluation can Distinguish the Good and Bad AI Responses to Patient Questions about Hospitalization
- Beyond Log Likelihood: Probability-Based Objectives for Supervised Fine-Tuning across the Model Capability Continuum
- RoBiologyDataChoiceQA: A Romanian Dataset for improving Biology understanding of Large Language Models
- The Dialogue That Heals: A Comprehensive Evaluation of Doctor Agents' Inquiry Capability
- KnowGuard: Knowledge-Driven Abstention for Multi-Round Clinical Reasoning
- AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification
- Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
- MedMMV: A Controllable Multimodal Multi-Agent Framework for Reliable and Verifiable Clinical Reasoning
- MDD-Thinker: Towards Large Reasoning Models for Major Depressive Disorder Diagnosis
- Beyond Overall Accuracy: A Psychometric Deep Dive into the Topic-Specific Medical Capabilities of 80 Large Language Models
- SynthPert: Enhancing LLM Biological Reasoning via Synthetic Reasoning Traces for Cellular Perturbation Prediction
- Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm
- Beyond the Exploration-Exploitation Trade-off: A Hidden State Approach for LLM Reasoning in RLVR
- Anchored Supervised Fine-Tuning
- Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMs
- No Loss, No Gain: Gated Refinement and Adaptive Compression for Prompt Optimization
- Model Consistency as a Cheap yet Predictive Proxy for LLM Elo Scores
- Evaluating Uncertainty Quantification Methods in Argumentative Large Language Models
- How Accurate Are LLMs at Multi-Question Answering on Conversational Transcripts?
- Correct Reasoning Paths Visit Shared Decision Pivots
- Eigen-1: Adaptive Multi-Agent Refinement with Monitor-Based RAG for Scientific Reasoning
- Integrated Framework for LLM Evaluation with Answer Generation
- RAR2: Retrieval-Augmented Medical Reasoning via Thought-Driven Retrieval
- Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generation
- Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models
- MedLLM: An Open Medical Language Model at the Sub-Billion Scale
- Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
- Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
- Filling in the Clinical Gaps in Benchmark: Case for HealthBench for the Japanese medical system
- Exploiting Tree Structure for Credit Assignment in RL Training of LLMs
- MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents
- Brittleness and Promise: Knowledge Graph Based Reward Modeling for Diagnostic Reasoning
- GRIL: Knowledge Graph Retrieval-Integrated Learning with Large Language Models
- From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
- RephQA: Evaluating Readability of Large Language Models in Public Health Question Answering
- Reward Hacking Mitigation using Verifiable Composite Rewards
- A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-making
- Enhancing Retrieval Augmentation via Adversarial Collaboration
- Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning
- Leveraging Large Language Models to Effectively Generate Visual Data for Canine Musculoskeletal Diagnoses
- MatQnA: A Benchmark Dataset for Multi-modal Large Language Models in Materials Characterization and Analysis
- Patient-Zero: Scaling Synthetic Patient Agents to Real-World Distributions without Real Patient Data
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
- Memorization in Large Language Models in Medicine: Prevalence, Characteristics, and Implications
- Transparency of medical artificial intelligence systems
- Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruning
- Performance Assessment Strategies for Generative AI Applications in Healthcare
- BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
- Does This Look Familiar to You? Knowledge Analysis via Model Internal Representations
- MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations
- The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
- Delta Activations: A Representation for Finetuned Large Language Models
- Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation
- PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
- MedQARo: A Large-Scale Benchmark for Medical Question Answering in Romanian
- JudgeAgent: Knowledge-wise and Dynamic LLM Evaluation with Agent-as-Interviewer
- MedOmni-45°: A Safety-Performance Benchmark for Reasoning-Oriented LLMs in Medicine
- RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
- OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries
- Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models
- A Graph-Based Test-Harness for LLM Evaluation
- DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
- Benchmarking GPT-5 for biomedical natural language processing
- AraHealthQA 2025: The First Shared Task on Arabic Health Question Answering
- CataractSurg-80K: Knowledge-Driven Benchmarking for Structured Reasoning in Ophthalmic Surgery Planning
- Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
- MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation
- Select to Know: An Internal-External Knowledge Self-Selection Framework for Domain-Specific Question Answering
- MedCoT-RAG: Causal Chain-of-Thought RAG for Medical Question Answering
- Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
- ProMed: Shapley Information Gain Guided Reinforcement Learning for Proactive Medical LLMs
- HeteroRAG: A Heterogeneous Retrieval-Augmented Generation Framework for Medical Vision Language Tasks
- SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression
- Exploring Efficiency Frontiers of Thinking Budget in Medical Reasoning: Scaling Laws between Computational Resources and Reasoning Quality
- QuarkMed Medical Foundation Model Technical Report
- Reverse Physician-AI Relationship: Full-process Clinical Diagnosis Driven by a Large Language Model
- Efficient Forward-Only Data Valuation for Pretrained LLMs and VLMs
- Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
- Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
- MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation
- Capabilities of GPT-5 on Multimodal Medical Reasoning
- TeamMedAgents: Pareto-Efficient Multi-Agent Medical Reasoning Through Teamwork Theory
- A Multi-Agent Approach to Neurological Clinical Reasoning
- Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond
- HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways
- UR2: Unify RAG and Reasoning through Reinforcement Learning
- Retrieval Augmented Large Language Model System for Comprehensive Drug Contraindications
- RTTC: Reward-Guided Collaborative Test-Time Compute
- Iterative Learning of Computable Phenotypes for Treatment Resistant Hypertension using Large Language Models
- PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment
- FedCoT: Communication-Efficient Federated Reasoning Enhancement for Large Language Models
- TRAIL: Joint Inference and Refinement of Knowledge Graphs with Large Language Models
- Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
- RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior
- MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
- MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
- HealthFlow: A Self-Evolving AI Agent with Meta Planning for Autonomous Healthcare Research
- Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data
- Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications
- Doctor Sun: A Bilingual Multimodal Large Language Model for Biomedical AI
- ControlMed: Adding Reasoning Control to Medical Language Model
- DeepSieve: Information Sieving via LLM-as-a-Knowledge-Router
- HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
- MediQAl: A French Medical Question Answering Dataset for Knowledge and Reasoning Evaluation
- Health Insurance Coverage Rule Interpretation Corpus: Law, Policy, and Medical Guidance for Health Insurance Coverage Understanding
- ChiMed 2.0: Advancing Chinese Medical Dataset in Facilitating Large Language Modeling
- From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation
- Toward Revealing Nuanced Biases in Medical LLMs
- Towards Domain Specification of Embedding Models in Medicine
- Adaptive Cluster Collaborativeness Boosts LLMs Medical Decision Support Capacity
- HIVMedQA: Benchmarking large language models for HIV medical decision support
- LoRA-Leak: Membership Inference Attacks Against LoRA Fine-tuned Language Models
- Critique of impure reason: Unveiling the reasoning behaviour of medical large language models
- Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
- Shaping capabilities with token-level data filtering
- MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
- TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
- WakenLLM: Evaluating Reasoning Potential and Stability in LLMs via Fine-Grained Benchmarking
- ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
- Med-REFL: Medical Reasoning Enhancement via Self-Corrected Fine-grained Reflection
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering
- Trustworthy AI for Medicine: Continuous Hallucination Detection and Elimination with CHECK
- Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
- Evaluating LLMs in Medicine: A Call for Rigor, Transparency
- Automating Expert-Level Medical Reasoning Evaluation of Large Language Models
- TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text
- IyawoBench v2.0: Extended Diagnostic Evaluation of Large Language Model Clinical Triage in Nigerian Primary Care
- Empowering Healthcare Practitioners with Language Models: Structuring Speech Transcripts in Two Real-World Clinical Applications
- Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
- BYOKG-RAG: Multi-Strategy Graph Retrieval for Knowledge Graph Question Answering
- Conversation Forests: The Key to Fine Tuning Large Language Models for Multi-Turn Medical Conversations is Branching
- Read Quietly, Think Aloud: Decoupling Comprehension and Reasoning in LLMs
- Conformal Information Pursuit for Interactively Guiding Large Language Models
- SynapseRoute: An Auto-Route Switching Framework on Dual-State Large Language Model
- DynamiCare: A Dynamic Multi-Agent Framework for Interactive and Open-Ended Medical Decision-Making
- CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics
- GeoRA: Geometry-Aware Low-Rank Adaptation for RLVR
- MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
- Enterprise Large Language Model Evaluation Benchmark
- Accurate and Energy Efficient: Local Retrieval-Augmented Generation Models Outperform Commercial Large Language Models in Medical Tasks
- LLM-Driven Medical Document Analysis: Enhancing Trustworthy Pathology and Differential Diagnosis
- MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration
- KokushiMD-10: Benchmark for Evaluating Large Language Models on Ten Japanese National Healthcare Licensing Examinations
- Enabling On-Device Medical AI Assistants via Input-Driven Saliency Adaptation
- EQuARX: Efficient Quantized AllReduce in XLA for Distributed Machine Learning Acceleration
- KAG-Thinker: Interactive Thinking and Deep Reasoning in LLMs via Knowledge-Augmented Generation
- From Raw Corpora to Domain Benchmarks: Automated Evaluation of LLM Domain Expertise
- Gazal-R1: Achieving State-of-the-Art Medical Reasoning with Parameter-Efficient Two-Stage Training
- Essential-Web v1.0: 24T tokens of organized web data
- Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
- Confident-Knowledge Diversity Drives Human-Human and Human-AI Free Discussion Synergy and Reveals Pure-AI Discussion Shortfalls
- Tiered Agentic Oversight: A Hierarchical Multi-Agent System for Healthcare Safety
- Training-free LLM Merging for Multi-task Learning
- Cascaded Language Models for Cost-effective Human-AI Decision-Making
- RAG+: Enhancing Retrieval-Augmented Generation with Application-Aware Reasoning
- Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards
- MedCite: Can Language Models Generate Verifiable Text for Medicine?
- BioClinical ModernBERT: A State-of-the-Art Long-Context Encoder for Biomedical and Clinical NLP
- Mind the Gap: Benchmarking LLM Uncertainty and Calibration with Specialty-Aware Clinical QA and Reasoning-Based Behavioural Features
- Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning
- C-PATH: Conversational Patient Assistance and Triage in Healthcare System
- Building Models of Neurological Language
- MIRIAD: Augmenting LLMs with millions of medical query-response pairs
- SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat
- Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
- Trustworthy Medical Question Answering: An Evaluation-Centric Survey
- MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science
- ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model Capabilities
- FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models
- Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
- PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs
- VM14K: First Vietnamese Medical Benchmark
- Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
- MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine
- Enhancing Clinical Multiple-Choice Questions Benchmarks with Knowledge Graph Guided Distractor Generation
- PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark
- Semi-structured LLM Reasoners Can Be Rigorously Audited
- Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning
- ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
- TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine
- MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering
- Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
- X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
- MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks
- Elicit and Enhance: Advancing Multimodal Reasoning in Medical Scenarios
- Second Opinion Matters: Towards Adaptive Clinical AI via the Consensus of Expert Model Ensemble
- CDR-Agent: Intelligent Selection and Execution of Clinical Decision Rules Using Large Language Model Agents
- ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room
- Resolving Knowledge Conflicts in Domain-specific Data Selection: A Case Study on Medical Instruction-tuning
- BioHopR: A Benchmark for Multi-Hop, Multi-Answer Reasoning in Biomedical Domain
- Reason-Align-Respond: Aligning LLM Reasoning with Knowledge Graphs for KGQA
- BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum
- MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems
- Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations
- Token Distillation: Attention-aware Input Embeddings For New Tokens
- Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and Opportunities
- DoctorRAG: Medical RAG Fusing Knowledge with Patient Analogy through Textual Gradients
- Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
- FairMedQA: Benchmarking Bias in Large Language Models for Medical Question Answering
- The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants
- MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
- Towards Harmonized Uncertainty Estimation for Large Language Models
- TAGS: A Test-Time Generalist-Specialist Framework with Retrieval-Augmented Reasoning and Verification
- Shared Prefixes, Better Credit: Adaptive Routing for Multi-Agent Reasoning
- Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL
- PatientSim: A Persona-Driven Simulator for Realistic Doctor-Patient Interactions
- WiNGPT-3.0 Technical Report
- PerMedCQA: Benchmarking Large Language Models on Medical Consumer Question Answering in Persian Language
- A Japanese Language Model and Three New Evaluation Benchmarks for Pharmaceutical NLP
- INFERENCEDYNAMICS: Efficient Routing Across LLMs through Structured Capability and Knowledge Profiling
- X-MAS: Towards Building Multi-Agent Systems with Heterogeneous LLMs
- Diagnosing our datasets: How does my language model learn clinical information?
- KaFT: Knowledge-aware Fine-tuning for Boosting LLMs' Domain-specific Question-Answering Performance
- sudoLLM: On Multi-role Alignment of Language Models
- DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
- s3: You Don't Need That Much Data to Train a Search Agent via RL
- MedBrowseComp: Benchmarking Medical Deep Research and Computer Use
- Context-Free Synthetic Data Mitigates Forgetting
- R2MED: A Benchmark for Reasoning-Driven Medical Retrieval
- HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding
- Decoding Rarity: Large Language Models in the Diagnosis of Rare Diseases
- Synthetic Data RL: Task Definition Is All You Need
- ExpertSteer: Intervening in LLMs through Expert Knowledge
- IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
- MoL for LLMs: Dual-Loss Optimization to Enhance Domain Expertise While Preserving General Capabilities
- AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation
- RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation
- MedCaseReasoning: Evaluating and learning diagnostic reasoning from clinical case reports
- Disentangling Reasoning and Knowledge in Medical Large Language Models
- A Systematic Analysis of Base Model Choice for Reward Modeling
- A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
- FactsR: A Safer Method for Producing High Quality Healthcare Documentation
- Large Language Models for Cancer Communication: Evaluating Linguistic Quality, Safety, and Accessibility in Generative AI
- Analog Foundation Models
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
- Large Language Models for Computer-Aided Design: A Survey
- C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning
- Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information
- Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies
- Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models
- Quantifying Hallucinations in Language Language Models on Medical Textbooks
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- Emotions in the Loop: A Survey of Affective Computing for Emotional Support
- Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications
- FLARE: Few-shot Learning-based Adaptive Reflective Engine
- On the Utility and Factual Reliability of Pruned Mixture-of-Experts Models in the Biomedical Domain
- Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination
- IntrAgent: An LLM Agent for Content-Grounded Information Retrieval through Literature Review
- Data Darwinism Part II: DataEvolve -- AI can Autonomously Evolve Pretraining Data Curation
- OpenBioRQ: Unsolved Biomedical Research Questions for Agents
- LLMs can construct powerful representations and streamline sample-efficient supervised learning
- A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design
- Flaws in the LLM Automation Narrative
- Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese
- Architecting Trust in Artificial Epistemic Agents
- OptiLeak: Efficient Prompt Reconstruction via Reinforcement Learning in Multi-tenant LLM Services
- Fully Open Meditron: An Auditable Pipeline for Clinical LLMs
- Symphony-Coord: Adaptive Routing for Multi-Agent LLM Systems
- SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
- Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning
- Generalistic or Specific Embeddings, Which is Better? An Empirical Study on Search for Clinical Coding in Non-English Languages
- MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional
- Talk Before You Retrieve: Agent-Led Discussions for Better RAG in Medical QA
- Multimodal Large Language Models for Medicine: A Comprehensive Survey
- Tripartite-GraphRAG via Plugin Ontologies
- BRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text
- When are likely answers right? On Sequence Probability and Correctness in LLMs
- PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective
- PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
- A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage
- Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization
- Process Reward Agents for Steering Knowledge-Intensive Reasoning
- MedGemma 1.5 Technical Report
- Counterfactual Cultural Cues Reduce Medical QA Accuracy in LLMs: Identifier vs Context Effects
- SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding
- FormationEval, an open multiple-choice benchmark for petroleum geoscience
- Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks
- DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning
- DeepInvert: Semi-Supervised Embedding Inversion Against Obfuscated Language Models
- When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs
- FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables
- Exploring How LLMs Capture and Represent Domain-Specific Knowledge
- Instruction-Tuning Data Synthesis from Scratch via Web Reconstruction
- A Case Study Exploring the Current Landscape of Synthetic Medical Record Generation with Commercial LLMs
- How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
- CDF-RAG: Causal Dynamic Feedback for Adaptive Retrieval-Augmented Generation
- Gauging Overprecision in LLMs: An Empirical Study
- FOCUS: Decoupling Expert Personas in LLMs to Enhance Domain Expert Capabilities
- MACRO: Markov Chain Routing of Transformer Layers
- CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
- Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
- CliniChat: A Multi-Source Knowledge-Driven Framework for Clinical Interview Dialogue Reconstruction and Evaluation
- DICE: A Framework for Dimensional and Contextual Evaluation of Language Models
- QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model
- MedHal: An Evaluation Dataset for Medical Hallucination Detection
- Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
- Graph-based Approaches and Functionalities in Retrieval-Augmented Generation: A Comprehensive Survey
Discussions
Related