TinyBERT: Distilling BERT for Natural Language Understanding
2019/09/23 by Xiaoqi Jiao, Yichun Yin, Jiao, Xiaoqi +13 · 155 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.1909.10351
Findings of EMNLP 2020; results have been updated; code and model: https://github.com/huawei-noah/Pretrained-Language-Model/tree/master/TinyBERT
openalex publication_date 2019/09/23 · arxiv created 2020/10/16 · arxiv updated 2020/10/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks. However, pre-trained language models are usually computationally expensive, so it is difficult to efficiently execute them on resource-restricted devices. To accelerate inference and reduce model size while maintaining accuracy, we first propose a novel Transformer distillation method that is specially designed for knowledge distillation (KD) of the Transformer-based models. By leveraging this new KD method, the plenty of knowledge encoded in a large teacher BERT can be effectively transferred to a small student Tiny-BERT. Then, we introduce a new two-stage learning framework for TinyBERT, which performs Transformer distillation at both the pretraining and task-specific learning stages. This framework ensures that TinyBERT can capture he general-domain as well as the task-specific knowledge in BERT. TinyBERT with 4 layers is empirically effective and achieves more than 96.8% the performance of its teacher BERTBASE on GLUE benchmark, while being 7.5x smaller and 9.4x faster on inference. TinyBERT with 4 layers is also significantly better than 4-layer state-of-the-art baselines on BERT distillation, with only about 28% parameters and about 31% inference time of them. Moreover, TinyBERT with 6 layers performs on-par with its teacher BERTBASE.
Citations
Cited by
- LLMCache: Layer-Wise Caching Strategies for Accelerated Reuse in Transformer Inference
- EdgeFlex-Transformer: Transformer Inference for Edge Devices
- TiME: Tiny Monolingual Encoders for Efficient NLP Pipelines
- Beyond Real Weights: Hypercomplex Representations for Stable Quantization
- Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment
- Odin: Oriented Dual-module Integration for Text-rich Network Representation Learning
- Towards Edge General Intelligence: Knowledge Distillation for Mobile Agentic AI
- Deterministic Continuous Replacement: Fast and Stable Module Replacement in Pretrained Transformers
- A Systematic Study of Compression Ordering for Large Language Models
- NX-CGRA: A Programmable Hardware Accelerator for Core Transformer Algorithms on Edge Devices
- When Structure Doesn't Help: LLMs Do Not Read Text-Attributed Graphs as Effectively as We Expected
- Unifying points of interest taxonomies: mapping OpenStreetMap tags to the Foursquare category system
- Dynamic Temperature Scheduler for Knowledge Distillation
- Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning
- Black-Box On-Policy Distillation of Large Language Models
- CleverBirds: A Multiple-Choice Benchmark for Fine-grained Human Knowledge Tracing
- Self-Correction Distillation for Structured Data Question Answering
- Two Heads are Better than One: Distilling Large Language Model Features Into Small Models with Feature Decomposition and Mixture
- MobileLLM-Pro Technical Report
- CAMP-HiVe: Cyclic Pair Merging based Efficient DNN Pruning with Hessian-Vector Approximation for Resource-Constrained Systems
- A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
- Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
- Reviving Stale Updates: Data-Free Knowledge Distillation for Asynchronous Federated Learning
- Elastic Architecture Search for Efficient Language Models
- Distilling Multilingual Vision-Language Models: When Smaller Models Stay Multilingual
- SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
- FakeZero: Real-Time, Privacy-Preserving Misinformation Detection for Facebook and X
- FHECore: Rethinking GPU Microarchitecture for Fully Homomorphic Encryption
- SwiftEmbed: Ultra-Fast Text Embeddings via Static Token Lookup for Real-Time Applications
- FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
- Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
- SindBERT, the Sailor: Charting the Seas of Turkish NLP
- TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
- Amplifying Prominent Representations in Multimodal Learning via Variational Dirichlet Process
- Mixture of Experts Approaches in Dense Retrieval Tasks
- Efficient Adaptive Transformer: An Empirical Study and Reproducible Framework
- CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
- A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- GUIDE: Guided Initialization and Distillation of Embeddings
- Downsized and Compromised?: Assessing the Faithfulness of Model Compression
- Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments
- Layer-wise dynamic rank for compressing large language models
- CURA: Size Isnt All You Need -- A Compact Universal Architecture for On-Device Intelligence
- Knowledge distillation through geometry-aware representational alignment
- RestoRect: Degraded Image Restoration via Latent Rectified Flow & Feature Distillation
- Progressive Weight Loading: Accelerating Initial Inference and Gradually Boosting Performance on Resource-Constrained Environments
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- MonoCon: A general framework for learning ultra-compact high-fidelity representations using monotonicity constraints
- Otters: An Energy-Efficient SpikingTransformer via Optical Time-to-First-Spike Encoding
- Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking
- LLM on a Budget: Active Knowledge Distillation for Efficient Classification of Large Text Corpora
- LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
- Spatio-Temporal Pruning for Compressed Spiking Large Language Models
- A Transformer-Based Cross-Platform Analysis of Public Discourse on the 15-Minute City Paradigm
- AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models
- Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation
- NoteBar: An AI-Assisted Note-Taking System for Personal Knowledge Management
- A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension
- Expandable Residual Approximation for Knowledge Distillation
- Efficient Large Language Models with Zero-Shot Adjustable Acceleration
- SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version
- An Empirical Study of Knowledge Distillation for Code Understanding Tasks
- Checkmate: interpretable and explainable RSVQA is the endgame
- Recent Advances in Transformer and Large Language Models for UAV Applications
- Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
- Personalized Product Search Ranking: A Multi-Task Learning Approach with Tabular and Non-Tabular Data
- Bhav-Net: Knowledge Transfer for Cross-Lingual Antonym vs Synonym Distinction via Dual-Space Graph Transformers
- GeRe: Towards Efficient Anti-Forgetting in Continual Learning of LLM via General Samples Replay
- Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code
- Distillation-Enhanced Clustering Acceleration for Encrypted Traffic Classification
- DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large Language Models
- Enhanced Arabic Text Retrieval with Attentive Relevance Scoring
- On the Sustainability of AI Inferences in the Edge
- Model-free Speculative Decoding for Transformer-based ASR with Token Map Drafting
- Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
- The Carbon Cost of Conversation, Sustainability in the Age of Language Models
- Collaborative Distillation Strategies for Parameter-Efficient Language Model Deployment
- Basic Reading Distillation
- Breaking the Tokenizer Barrier: On-Policy Distillation across Model Families
- OPRD: On-Policy Representation Distillation
- DyG-RAG: Dynamic Graph Retrieval-Augmented Generation with Event-Centric Reasoning
- Seq vs Seq: An Open Suite of Paired Encoders and Decoders
- DQLoRA: A Lightweight Domain-Aware Denoising ASR via Adapter-guided Distillation
- L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training
- Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks
- Put Teacher in Student's Shoes: Cross-Distillation for Ultra-compact Model Compression Framework
- Tractable Representation Learning with Probabilistic Circuits
- QFFN-BERT: An Empirical Study of Depth, Performance, and Data Efficiency in Hybrid Quantum-Classical Transformers
- MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants
- Matterhorn: Masked Time-to-First-Spike Encoding by Reassigning the Silent State for Sparse and Energy-Efficient Spiking Transformers
- Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
- MobileRAG: A Fast, Memory-Efficient, and Energy-Efficient Method for On-Device RAG
- Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models
- Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation
- Scaling Transformers for Time Series Forecasting: Do Pretrained Large Models Outperform Small-Scale Alternatives?
- Multimodal Medical Image Binding via Shared Text Embeddings
- From Tiny Machine Learning to Tiny Deep Learning: A Survey
- A Hybrid DeBERTa and Gated Broad Learning System for Cyberbullying Detection in English Text
- Knowledge Distillation Framework for Accelerating High-Accuracy Neural Network-Based Molecular Dynamics Simulations
- AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes
- Pre-trained Summarization Distillation
- Learning Student-Friendly Teacher Networks for Knowledge Distillation
- A Heuristic Perspective on Debiasing Language Models
- Analyzing Transformer Models and Knowledge Distillation Approaches for Image Captioning on Edge AI
- Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning
- On Fairness of Task Arithmetic: The Role of Task Vectors
- RCCDA: Adaptive Model Updates in the Presence of Concept Drift under a Constrained Resource Budget
- FLAT-LLM: Fine-grained Low-rank Activation Space Transformation for Large Language Model Compression
- Efficient Large Language Model Inference with Neural Block Linearization
- Real-Time Execution of Large-scale Language Models on Mobile
- Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation
- SEMFED: Semantic-Aware Resource-Efficient Federated Learning for Heterogeneous NLP Tasks
- Avoid Forgetting by Preserving Global Knowledge Gradients in Federated Learning with Non-IID Data
- Online Knowledge Distillation with Reward Guidance
- Language Model Distillation: A Temporal Difference Imitation Learning Perspective
- FAR: Function-preserving Attention Replacement for IMC-friendly Inference
- On the creation of narrow AI: hierarchy and nonlocality of neural network skills
- InfiGFusion: Graph-on-Logits Distillation via Efficient Gromov-Wasserstein for Model Fusion
- Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization
- Latent Flow Transformer
- Quantum Knowledge Distillation for Large Language Models
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
- Bridging Generative and Discriminative Learning: Few-Shot Relation Extraction via Two-Stage Knowledge-Guided Pre-training
- ExpertSteer: Intervening in LLMs through Expert Knowledge
- On Membership Inference Attacks in Knowledge Distillation
- Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge Distillation
- Tracr-Injection: Distilling Algorithms into Pre-trained Language Models
- Private Transformer Inference in MLaaS: A Survey
- Progressive2: A Teacher-Student Progressive Co-Evolving Knowledge Distillation Method for Substantial Model Compression
- PWC-MoE: Privacy-Aware Wireless Collaborative Mixture of Experts
- Prefix-Guided On-Policy Distillation: Mining Golden Trajectories from Rollouts
- TURL: Table Understanding through Representation Learning
- KDH-MLTC: Knowledge Distillation for Healthcare Multi-Label Text Classification
- Efficient Fine-Tuning of Quantized Models via Adaptive Rank and Bitwidth
- Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
- Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?
- LLM-Based Threat Detection and Prevention Framework for IoT Ecosystems
- SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extraction
- TIDAL: Recovering Temporal Phase for Cloud Block Storage Placement from LLM-Derived Semantics
- Kernel Affine Hull Machines as Compute-Efficient Encoders for Frozen Semantic Spaces
- Weak-to-Strong Generalization via Direct On-Policy Distillation
- Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
- Mixture of Experts for Decentralized Generative AI and Reinforcement Learning in Wireless Networks: A Comprehensive Survey
- KETCHUP: K-Step Return Estimation for Sequential Knowledge Distillation
- LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
- LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
- HMI: Hierarchical Knowledge Management for Efficient Multi-Tenant Inference in Pretrained Language Models
- A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms
- Saliency-driven Dynamic Token Pruning for Large Language Models
- Honey, I Shrunk the Language Model: Impact of Knowledge Distillation Methods on Performance and Explainability
- W-PCA Based Gradient-Free Proxy for Efficient Search of Lightweight Language Models
- DistilQwen2.5: Industrial Practices of Training Distilled Open Lightweight Language Models
- Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
- A Dual-Space Framework for General Knowledge Distillation of Large Language Models
- Multi-Sense Embeddings for Language Models and Knowledge Distillation
Related