Knowledge Distillation: A Survey
2021/03/22 by Jianping Gou, Baosheng Yu, Stephen J. Maybank +1 · 141 citations
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications
paper · doi:10.1007/s11263-021-01453-z
openalex publication_date 2021/03/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Cited by
- UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
- BTKD++: Beyond Teachers by Critically Distilling Knowledge from Teacher’s Bias
- DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA
- EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens
- Towards Edge General Intelligence: Knowledge Distillation for Mobile Agentic AI
- DualGazeNet: A Biologically Inspired Dual-Gaze Query Network for Salient Object Detection
- Towards Characterizing Knowledge Distillation of PPG Heart Rate Estimation Models
- CoS: Towards Optimal Event Scheduling via Chain-of-Scheduling
- BD-Net: Has Depth-Wise Convolution Ever Been Applied in Binary Neural Networks?
- Uncertainty-Resilient Multimodal Learning via Consistency-Guided Cross-Modal Transfer
- Generalizable and Efficient Automated Scoring with a Knowledge-Distilled Multi-Task Mixture-of-Experts
- DINO-MX: A Modular & Flexible Framework for Self-Supervised Learning
- D2-VPR: A Parameter-efficient Visual-foundation-model-based Visual Place Recognition Method via Knowledge Distillation and Deformable Aggregation
- Fast Reasoning Segmentation for Images and Videos
- Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
- BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning
- EarthSight: A Distributed Framework for Low-Latency Satellite Intelligence
- ConSurv: Multimodal Continual Learning for Survival Analysis
- Two Heads are Better than One: Distilling Large Language Model Features Into Small Models with Feature Decomposition and Mixture
- Practical Policy Distillation for Reinforcement Learning in Radio Access Networks
- DWM-RO: Decentralized World Models with Reasoning Offloading for SWIPT-enabled Satellite-Terrestrial HetNets
- TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
- Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
- Democratizing LLM Efficiency: From Hyperscale Optimizations to Universal Deployability
- Privacy-Aware Continual Self-Supervised Learning on Multi-Window Chest Computed Tomography for Domain-Shift Robustness
- Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base
- maxVSTAR: Maximally Adaptive Vision-Guided CSI Sensing with Closed-Loop Edge Model Adaptation for Robust Human Activity Recognition
- Deep Long-Tailed Learning: A Survey
- Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment
- A Survey of Imbalanced Learning on Graphs: Problems, Techniques, and Future Directions
- Fine-tuning Small Language Models (SLMs) for autonomous web-based geographical information systems (AWebGIS)
- SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
- Optimizing Retrieval for RAG via Reinforced Contrastive Learning
- UHKD: A Unified Framework for Heterogeneous Knowledge Distillation via Frequency-Domain Representations
- HYDRA: HYbrid knowledge Distillation and spectral Reconstruction Algorithm for high channel hyperspectral camera applications
- Hankel Singular Value Regularization for Highly Compressible State Space Models
- Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- Transforming volcanic monitoring: A dataset and benchmark for onboard volcano activity detection
- Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
- Compressing Quaternion Convolutional Neural Networks for Audio Classification
- TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
- KCM: KAN-Based Collaboration Models Enhance Pretrained Large Models
- MobiAct: Efficient MAV Action Recognition Using MobileNetV4 with Contrastive Learning and Knowledge Distillation
- Learning Task-Agnostic Representations through Multi-Teacher Distillation
- A Rectification-Based Approach for Distilling Boosted Trees into Decision Trees
- Black-Box Evasion Attacks on Data-Driven Open RAN Apps: Tailored Design and Experimental Evaluation
- Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
- MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
- Faster Molecular Dynamics with Neural Network Potentials via Distilled Multiple Time-Stepping and Nonconservative Forces
- TeamFormer: Shallow Parallel Transformers with Progressive Approximation
- Spiking Neural Network Architecture Search: A Survey
- Information-Theoretic Criteria for Knowledge Distillation in Multimodal Learning
- Information-Preserving Reformulation of Reasoning Traces for Antidistillation
- BioX-Bridge: Model Bridging for Unsupervised Cross-Modal Knowledge Transfer across Biosignals
- Optimally Deep Networks -- Adapting Model Depth to Datasets for Superior Efficiency
- SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
- Defense against Unauthorized Distillation in Image Restoration via Feature Space Perturbation
- Optimizing delivery for quick commerce factoring qualitative assessment of generated routes
- GCPO: When Contrast Fails, Go Gold
- LiveThinking: Enabling Real-Time Efficient Reasoning for AI-Powered Livestreaming via Reinforcement Learning
- Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
- Recover-LoRA: Data-Free Accuracy Recovery of Degraded Language Models via Low-Rank Adaptation
- COLE: a Comprehensive Benchmark for French Language Understanding Evaluation
- Distilling Reasoning into Student LLMs: Local Naturalness for Selecting Teacher Data
- SDAKD: Student Discriminator Assisted Knowledge Distillation for Super-Resolution Generative Adversarial Networks
- MECKD: Deep Learning-Based Fall Detection in Multilayer Mobile Edge Computing With Knowledge Distillation
- TAP: Two-Stage Adaptive Personalization of Multi-task and Multi-Modal Foundation Models in Federated Learning
- Circuit Distillation
- Federated Learning Meets LLMs: Feature Extraction From Heterogeneous Clients
- UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
- A Fast Knowledge Distillation Framework for Visual Recognition
- Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack
- An Encoder-Decoder Network for Beamforming over Sparse Large-Scale MIMO Channels
- PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
- Streamline pathology foundation model by cross-magnification distillation
- Progressive Weight Loading: Accelerating Initial Inference and Gradually Boosting Performance on Resource-Constrained Environments
- Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
- POEM: Explore Unexplored Reliable Samples to Enhance Test-Time Adaptation
- LogReasoner: Empowering LLMs with Expert-like Coarse-to-Fine Reasoning for Automated Log Analysis
- Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
- MDE for crop representations in smart farming digital twins: a reinforcement learning perspective
- Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
- The First Cadenza Challenges: Using Machine Learning Competitions to Improve Music for Listeners With a Hearing Loss
- Dark3R: Learning Structure from Motion in the Dark
- The Hallmarks of Predictive Oncology
- PPG-Distill: Efficient Photoplethysmography Signals Analysis via Foundation Model Distillation
- Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models
- SOLAR: Switchable Output Layer for Accuracy and Robustness in Once-for-All Training
- A Unified AI Approach for Continuous Monitoring of Human Health and Diseases from Intensive Care Unit to Home with Physiological Foundation Models (UNIPHY+)
- MMCD: Multi-Modal Collaborative Decision-Making for Connected Autonomy with Knowledge Distillation
- Who to Trust? Aggregating Client Knowledge in Logit-Based Federated Learning
- Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
- Ensemble of Pre-Trained Models for Long-Tailed Trajectory Prediction
- iCD: A Implicit Clustering Distillation Mathod for Structural Information Mining
- Generative AI for Multimedia Communication: Recent Advances, An Information-Theoretic Framework, and Future Opportunities
- One Teacher is Enough? Pre-trained Language Model Distillation from Multiple Teachers
- Learning from Diverse Reasoning Paths with Routing and Collaboration
- Ensemble Distribution Distillation for Self-Supervised Human Activity Recognition
- Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
- MEGG: Replay via Maximally Extreme GGscore in Incremental Learning for Neural Recommendation Models
- Online Clustering of Seafloor Imagery for Interpretation during Long-Term AUV Operations
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- Toward a Unified Framework for Debugging Concept-based Models
- Comparative Analysis of Transformer Models in Disaster Tweet Classification for Public Safety
- Deep Self-knowledge Distillation: A hierarchical supervised learning for coronary artery segmentation
- Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling
- TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization
- UrbanInsight: A Distributed Edge Computing Framework with LLM-Powered Data Filtering for Smart City Digital Twins
- I Stolenly Swear That I Am Up to (No) Good: Design and Evaluation of Model Stealing Attacks
- Towards On-Device Personalization: Cloud-device Collaborative Data Augmentation for Efficient On-device Language Model
- PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
- Federated Learning for Large Models in Medical Imaging: A Comprehensive Review
- Dual-Model Weight Selection and Self-Knowledge Distillation for Medical Image Classification
- ATMS-KD: Adaptive Temperature and Mixed Sample Knowledge Distillation for a Lightweight Residual CNN in Agricultural Embedded Systems
- Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning
- Deep learning in automated ultrasonic NDE – Developments, axioms and opportunities
- Synthetic Adaptive Guided Embeddings (SAGE): A Novel Knowledge Distillation Method
- Neural Network Quantization for Microcontrollers: A Comprehensive Survey of Methods, Platforms, and Applications
- Seeing Further on the Shoulders of Giants: Knowledge Inheritance for Vision Foundation Models
- Data Distillation for Text Classification
- Whispering Context: Distilling Syntax and Semantics for Long Speech Transcripts
- Unified Knowledge Distillation Framework: Fine-Grained Alignment and Geometric Relationship Preservation for Deep Face Recognition
- LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
- HyperKD: Distilling Cross-Spectral Knowledge in Masked Autoencoders via Inverse Domain Shift with Spatial-Aware Masking and Specialized Loss
- Toward Robust Semi-supervised Regression via Dual-stream Knowledge Distillation
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Transferring Social Network Knowledge from Multiple GNN Teachers to Kolmogorov-Arnold Networks
- Data Exfiltration by Compression Attack: Definition and Evaluation on Medical Image Data
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- NT-ML: Backdoor Defense via Non-target Label Training and Mutual Learning
- MedMambaLite: Hardware-Aware Mamba for Medical Image Classification
- Fine-Tuning Small Language Models (SLMs) for Autonomous Web-based Geographical Information Systems (AWebGIS)
- Slice or the Whole Pie? Utility Control for AI Models
- NeuroSync: Intent-Aware Code-Based Problem Solving via Direct LLM Understanding Modification
- CORE-ReID: Comprehensive Optimization and Refinement through Ensemble fusion in Domain Adaptation for person re-identification
- Realizing Scaling Laws in Recommender Systems: A Foundation-Expert Paradigm for Hyperscale Model Deployment
- Adaptive Knowledge Distillation for Device-Directed Speech Detection
- Slot Attention with Re-Initialization and Self-Distillation
- Teaching the Teacher: Improving Neural Network Distillability for Symbolic Regression via Jacobian Regularization
- Resource-Efficient Automatic Software Vulnerability Assessment via Knowledge Distillation and Particle Swarm Optimization
- Teach Me to Trick: Exploring Adversarial Transferability via Knowledge Distillation