Advances in Large Language Models for Medicine
2025/09/23 by Zhiyu Kan, Kan, Zhiyu, Wensheng Gan +5
Computer Science · Medicine · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #FOS: Computer and information sciences #Machine Learning in Healthcare #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2509.18690
openalex publication_date 2025/09/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Artificial intelligence (AI) technology has advanced rapidly in recent years, with large language models (LLMs) emerging as a significant breakthrough. LLMs are increasingly making an impact across various industries, with the medical field standing out as the most prominent application area. This paper systematically reviews the up-to-date research progress of LLMs in the medical field, providing an in-depth analysis of training techniques for large medical models, their adaptation in healthcare settings, related applications, as well as their strengths and limitations. Furthermore, it innovatively categorizes medical LLMs into three distinct types based on their training methodologies and classifies their evaluation approaches into two categories. Finally, the study proposes solutions to existing challenges and outlines future research directions based on identified issues in the field of medical LLMs. By systematically reviewing previous and advanced research findings, we aim to highlight the necessity of developing medical LLMs, provide a deeper understanding of their current state of development, and offer clear guidance for subsequent research.
Citations
- AutoMedPrompt: A New Framework for Optimizing LLM Medical Prompts Using Textual Gradients
- Mixture of Experts (MoE): A Big Data Perspective
- Leveraging Large Language Models for Enhancing Autonomous Vehicle Perception
- Unified Generative and Discriminative Training for Multi-modal Large Language Models
- Adaptive Self-Supervised Learning Strategies for Dynamic On-Device LLM Personalization
- Understanding the Performance and Estimating the Cost of LLM Fine-Tuning
- Large Language Models for Medicine: A Survey
- Benchmarking Retrieval-Augmented Generation for Medicine
- Retrieval-Augmented Generation for Large Language Models: A Survey
- LLMs Accelerate Annotation for Medical Information Extraction
- Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine
- Large Language Models in Law: A Survey
- Multimodal Large Language Models: A Survey
- Large Language Models in Education: Vision and Opportunities
- Large Language Models for Robotics: A Survey
- Model-as-a-Service (MaaS): A Survey
- A Survey of Large Language Models in Medicine: Progress, Application, and Challenge
- Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare
- BianQue: Balancing the Questioning and Suggestion Ability of Health LLMs with Multi-turn Health Conversations Polished by ChatGPT
- CareerX: A Retrieval-Augmented Generation Framework for Personalized AI-Driven Career Guidance
- Qilin-Med: Multi-stage Knowledge Injection Advanced Medical Large Language Model
- A Survey of Large Language Models for Healthcare: from Data, Technology, and Applications to Accountability and Ethics
- AnglE-optimized Text Embeddings
- CPLLM: Clinical Prediction with Large Language Models
- Baichuan 2: Open Large-scale Language Models
- C-Pack: Packed Resources For General Chinese Embeddings
- A Survey of Hallucination in Large Foundation Models
- Zhongjing: Enhancing the Chinese Medical Capabilities of Large Language Model through Expert Feedback and Real-world Multi-turn Dialogue
- Med-Flamingo: a Multimodal Medical Few-shot Learner
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Large language models in medicine
- GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
- Exploring and Characterizing Large Language Models For Embedded System Development and Debugging
- ClinicalGPT: Large Language Models Finetuned with Diverse Medical Data and Comprehensive Evaluation
- From Large Language Models to Databases and Back: A discussion on research and education
- From Large Language Models to Databases and Back: A Discussion on Research and Education
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- HuatuoGPT, towards Taming Language Model to Be a Doctor
- LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models
- Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding
- MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data
- HuaTuo: Tuning LLaMA Model with Chinese Medical Knowledge
- ChatGPT: More than a Weapon of Mass Deception, Ethical challenges and responses from the Human-Centered Artificial Intelligence (HCAI) perspective
- ChatGPT: More Than a “Weapon of Mass Deception” Ethical Challenges and Responses from the Human-Centered Artificial Intelligence (HCAI) Perspective
- LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models
- DoctorGLM: Fine-tuning your Chinese Doctor is not a Herculean Task
- ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain Knowledge
- Frozen Language Model Helps ECG Zero-Shot Learning
- DeID-GPT: Zero-shot Medical Text De-Identification by GPT-4
- GPT-4 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- Language Models are Few-shot Learners for Prognostic Prediction
- Truth Machines: Synthesizing Veracity in AI Language Models
- Large Language Models Encode Clinical Knowledge
- Large language models encode clinical knowledge
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions
- Prompted Opinion Summarization with GPT-3.5
- Metaverse Security and Privacy: An Overview
- Exploiting Pretrained Biochemical Language Models for Targeted Drug Design
- PaLM: Scaling Language Modeling with Pathways
- Training language models to follow instructions with human feedback
- GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Anomaly Rule Detection in Sequence Data
- M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining
- Finetuned Language Models Are Zero-Shot Learners
- The Power of Scale for Parameter-Efficient Prompt Tuning
- GLM: General Language Model Pretraining with Autoregressive Blank Infilling
- Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing
- The need for empathetic healthcare systems
- Language Models are Few-Shot Learners
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- MedDialog: Two Large-scale Medical Dialogue Datasets
- An overview of clinical decision support systems: benefits, risks, and strategies for success
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Parameter-Efficient Transfer Learning for NLP
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Related