Large language models propagate race-based medicine
2023/10/20 by Jesutofunmi A. Omiye, Jenna Lester, Simon Spichak +2 · 1 voice · 41 citations
Computer Science · Health Professions · Medicine · #Artificial Intelligence in Healthcare and Education #Interpreting and Communication in Healthcare #Topic Modeling
paper · pdf · doi:10.1038/s41746-023-00939-z
openalex publication_date 2023/10/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
Large language models (LLMs) are being integrated into healthcare systems; but these models may recapitulate harmful, race-based medicine. The objective of this study is to assess whether four commercially available large language models (LLMs) propagate harmful, inaccurate, race-based content when responding to eight different scenarios that check for race-based medicine or widespread misconceptions around race. Questions were derived from discussions among four physician experts and prior work on race-based medical misconceptions believed by medical trainees. We assessed four large language models with nine different questions that were interrogated five times each with a total of 45 responses per model. All models had examples of perpetuating race-based medicine in their responses. Models were not always consistent in their responses when asked the same question repeatedly. LLMs are being proposed for use in the healthcare setting, with some models already connecting to electronic health record systems. However, this study shows that based on our findings, these LLMs could potentially cause harm by perpetuating debunked, racist ideas.
Citations
Cited by
- Large language models, social demography, and hegemony: comparing authorship in human and synthetic text
- Hearsay: Vision-Language Medical Diagnoses Without an Image
- Exploring the Role of Generative AI in Dementia Resilience Building Activities: Uncovering Opportunities and Challenges
- The patient-safety implications of AI-based communication with migrants in general practice: a scoping review
- Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
- How human–AI feedback loops alter human perceptual, emotional and social judgements
- Uncovering Overconfident Failures in CXR Models via Augmentation-Sensitivity Risk Scoring
- MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
- Quantifying uncert-AI-nty: Testing the accuracy of LLMs’ confidence judgments
- AgentClinic: a multimodal benchmark for tool-using clinical AI agents
- Assessing and alleviating state anxiety in large language models
- Equitable Artificial Intelligence in Obstetrics, Maternal–Fetal Medicine, and Neonatology
- Qualitative Research in an Era of AI: A Pragmatic Approach to Data Analysis, Workflow, and Computation
- Do Large Language Models Favor Recent Content? A Study on Recency Bias in LLM-Based Reranking
- Compartmentalised Agentic Reasoning for Clinical NLI
- Walking Backward to Ensure Risk Management of Large Language Models in Medicine
- Usability evaluation and reporting for mobile health apps targeting patients with skin diseases: a systematic review
- Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
- Differentiating hype from practical applications of large language models in medicine -- a primer for healthcare professionals
- HIVMedQA: Benchmarking large language models for HIV medical decision support
- "It looks sexy but it's wrong." Tensions in creativity and accuracy using genAI for biomedical visualization
- Artificial intelligence and perspective for rare genetic kidney diseases
- Dr. GPT Will See You Now, but Should It? Exploring the Benefits and Harms of Large Language Models in Medical Diagnosis using Crowdsourced Clinical Cases
- Trustworthy Medical Question Answering: An Evaluation-Centric Survey
- Impact of Exposure Parameters on Deep Learning Models in Chest Radiography and Implications for Deployment
- A Typology of Synthetic Datasets for Dialogue Processing in Clinical Contexts
- FairMedQA: Benchmarking Bias in Large Language Models for Medical Question Answering
- Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
- Testing and Evaluation of Health Care Applications of Large Language Models
- Detecting Prefix Bias in LLM-based Reward Models
- Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Models
- Benchmarking Ethical and Safety Risks of Healthcare LLMs in China-Toward Systemic Governance under Healthy China 2030
- Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI
- Interaction Configurations and Prompt Guidance in Conversational AI for Question Answering in Human-AI Teams
- Gender-Dependent Diagnostic Substitution in LLM Medical Triage: Same Symptoms, Unequal Urgency
- Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual Classifiers
- Generative AI Literacy: A Comprehensive Framework for Literacy and Responsible Use
- From Promising Capability to Pervasive Bias: Assessing Large Language Models for Emergency Department Triage
- Beyond Misinformation: A Conceptual Framework for Studying AI Hallucinations in (Science) Communication
- Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations
- Navigating the Rabbit Hole: Emergent Biases in LLM-Generated Attack Narratives Targeting Mental Health Groups
Discussions
Related