A Survey on Knowledge Distillation of Large Language Models
2024/02/20 by Xiaohan Xu, Xu, Xiaohan, Ming Li +15 · 2 voices · 106 citations
Chemistry · Computer Science · #Chemistry #Chromatography #Computer science #Distillation #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2402.13116
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/02/20 · arxiv published 2024/02/20 · arxiv updated 2024/10/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/02
Abstract
In the era of Large Language Models (LLMs), Knowledge Distillation (KD) emerges as a pivotal methodology for transferring advanced capabilities from leading proprietary LLMs, such as GPT-4, to their open-source counterparts like LLaMA and Mistral. Additionally, as open-source LLMs flourish, KD plays a crucial role in both compressing these models, and facilitating their self-improvement by employing themselves as teachers. This paper presents a comprehensive survey of KD's role within the realm of LLM, highlighting its critical function in imparting advanced knowledge to smaller models and its utility in model compression and self-improvement. Our survey is meticulously structured around three foundational pillars: algorithm, skill, and verticalization -- providing a comprehensive examination of KD mechanisms, the enhancement of specific cognitive abilities, and their practical implications across diverse fields. Crucially, the survey navigates the intricate interplay between data augmentation (DA) and KD, illustrating how DA emerges as a powerful paradigm within the KD framework to bolster LLMs' performance. By leveraging DA to generate context-rich, skill-specific training data, KD transcends traditional boundaries, enabling open-source models to approximate the contextual adeptness, ethical alignment, and deep semantic insights characteristic of their proprietary counterparts. This work aims to provide an insightful guide for researchers and practitioners, offering a detailed overview of current methodologies in KD and proposing future research directions. Importantly, we firmly advocate for compliance with the legal terms that regulate the use of LLMs, ensuring ethical and lawful application of KD of LLMs. An associated Github repository is available at https://github.com/Tebmer/Awesome-Knowledge-Distillation-of-LLMs.
Cited by
- HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
- GLaD: Geometric Latent Distillation for Vision-Language-Action Models
- MemLoRA: Distilling Expert Adapters for On-Device Memory Systems
- Colon-X: Advancing Intelligent Colonoscopy toward Clinical Reasoning
- TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
- Divide, Cache, Conquer: Dichotomic Prompting for Efficient Multi-Label LLM-Based Classification
- Reviving Stale Updates: Data-Free Knowledge Distillation for Asynchronous Federated Learning
- Improving LLM Reasoning via Dependency-Aware Query Decomposition and Logic-Parallel Content Expansion
- Beyond Neural Incompatibility: Easing Cross-Scale Knowledge Transfer in Large Language Models through Latent Semantic Alignment
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- The Structural Scalpel: Automated Contiguous Layer Pruning for Large Language Models
- Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
- CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
- Online In-Context Distillation for Low-Resource Vision Language Models
- Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
- Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests
- AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
- AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- BIRD-INTERACT: Re-imagining Text-to-SQL Evaluation for Large Language Models via Lens of Dynamic Interactions
- Circuit Distillation
- DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
- Reasoning Scaffolding: Distilling the Flow of Thought from LLMs
- Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
- Dual-View Alignment Learning with Hierarchical-Prompt for Class-Imbalance Multi-Label Classification
- Understanding Post-Training Structural Changes in Large Language Models
- MMCD: Multi-Modal Collaborative Decision-Making for Connected Autonomy with Knowledge Distillation
- SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
- LLM-Driven Policy Diffusion: Enhancing Generalization in Offline Reinforcement Learning
- Towards On-Device Personalization: Cloud-device Collaborative Data Augmentation for Efficient On-device Language Model
- Uncertainty-Aware Collaborative System of Large and Small Models for Multimodal Sentiment Analysis
- The point is the mask: scaling coral reef segmentation with weak supervision
- From Sound to Sight: Towards AI-authored Music Videos
- Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training
- GeRe: Towards Efficient Anti-Forgetting in Continual Learning of LLM via General Samples Replay
- Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code
- LMAR: Language Model Augmented Retriever for Domain-specific Knowledge Indexing
- PROVCREATOR: Synthesizing Complex Heterogenous Graphs with Node and Edge Attributes
- Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
- Collaborative Distillation Strategies for Parameter-Efficient Language Model Deployment
- ERNIE 5.0 Technical Report
- Synthetic Voice Data for Automatic Speech Recognition in African Languages
- Theory-Grounded Evaluation of Human-Like Fallacy Patterns in LLM Reasoning
- FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
- Towards Interpretable Time Series Foundation Models
- Exploiting Edge Features for Transferable Adversarial Attacks in Distributed Machine Learning
- Lyria: A Genetic Algorithm-Driven Neuro-Symbolic Reasoning Framework for LLMs
- ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
- Learning from Random Subspace Exploration: Generalized Test-Time Augmentation with Self-supervised Distillation
- Towards a Small Language Model Lifecycle Framework
- ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs
- Event-Priori-Based Vision-Language Model for Efficient Visual Understanding
- Biased Teacher, Balanced Student
- Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration
- CKD-EHR:Clinical Knowledge Distillation for Electronic Health Records
- Dataset distillation for memorized data: Soft labels can leak held-out teacher knowledge
- Bridging the Digital Divide: Small Language Models as a Pathway for Physics and Photonics Education in Underdeveloped Regions
- What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding
- A Layered Self-Supervised Knowledge Distillation Framework for Efficient Multimodal Learning on the Edge
- Being Strong Progressively! Enhancing Knowledge Distillation of Large Language Models through a Curriculum Learning Framework
- Optimization Problem Solving Can Transition to Evolutionary Agentic Workflows
- A Platform for Investigating Public Health Content with Efficient Concern Classification
- KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
- Common Inpainted Objects In-N-Out of Context
- Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition
- CAST: Contrastive Adaptation and Distillation for Semi-Supervised Instance Segmentation
- From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications
- ResSVD: Residual Compensated SVD for Large Language Model Compression
- Quantitative Analysis of Performance Drop in DeepSeek Model Quantization
- DOGe: Defensive Output Generation for LLM Protection Against Knowledge Distillation
- The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
- The Real Barrier to LLM Agent Usability is Agentic ROI
- Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN
- RAP: Runtime Adaptive Pruning for LLM Inference
- VERDI: VLM-Embedded Reasoning for Autonomous Driving
- Robust and Efficient AI-Based Attack Recovery in Autonomous Drones
- Exploring Federated Pruning for Large Language Models
- On Membership Inference Attacks in Knowledge Distillation
- Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge Distillation
- Dist2ill: Distributional Distillation for One-Pass Uncertainty Estimation in Large Language Models
- A Survey on Deep Neural Network Pruning: Taxonomy, Comparison, Analysis, and Recommendations
- TrialMatchAI: An End-to-End AI-powered Clinical Trial Recommendation System to Streamline Patient-to-Trial Matching
- A Survey on Collaborative Mechanisms Between Large and Small Language Models
- SOD: Step-wise On-policy Distillation for Small Language Model Agents
- ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
- Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications
- Antidistillation Fingerprinting
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- When Context Returns: Toward Robust Internalization in On-Policy Distillation
- Hierarchical Prompt-Domain Control and Learning for Resource-Constrained Agentic Language Models
- From Precision to Perception: User-Centred Evaluation of Keyword Extraction Algorithms for Internet-Scale Contextual Advertising
- Knowledge Distillation of Domain-adapted LLMs for Question-Answering in Telecom
- Towards Faster and More Compact Foundation Models for Molecular Property Prediction
- Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge
- Towards Harnessing the Collaborative Power of Large and Small Models for Domain Tasks
- Transforming remanufacturing automation with large language models: A forward-looking analysis with case studies
- Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion
- Honey, I Shrunk the Language Model: Impact of Knowledge Distillation Methods on Performance and Explainability
- Transferable Deployment of Semantic Edge Inference Systems via Unsupervised Domain Adaption
- The Effects of Grouped Structural Global Pruning of Vision Transformers on Domain Generalisation
- A Dual-Space Framework for General Knowledge Distillation of Large Language Models
- How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
- Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
- CAReDiO: Cultural Alignment via Representativeness and Distinctiveness Guided Data Optimization
- GAAPO: Genetic Algorithmic Applied to Prompt Optimization
- State Tuning: State-based Test-Time Scaling on RWKV-7
Discussions
Related