MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
2020/02/25 by Wenhui Wang, Furu Wei, Wang, Wenhui +9 · 210 citations
Computer Science · Engineering · Materials Science · #Topic Modeling #Ferroelectric and Negative Capacitance Devices #Machine Learning in Materials Science
paper · pdf · doi:10.48550/arxiv.2002.10957
Abstract
Pre-trained language models (e.g., BERT (Devlin et al., 2018) and its variants) have achieved remarkable success in varieties of NLP tasks. However, these models usually consist of hundreds of millions of parameters which brings challenges for fine-tuning and online serving in real-life applications due to latency and capacity constraints. In this work, we present a simple and effective approach to compress large Transformer (Vaswani et al., 2017) based pre-trained models, termed as deep self-attention distillation. The small model (student) is trained by deeply mimicking the self-attention module, which plays a vital role in Transformer networks, of the large model (teacher). Specifically, we propose distilling the self-attention module of the last Transformer layer of the teacher, which is effective and flexible for the student. Furthermore, we introduce the scaled dot-product between values in the self-attention module as the new deep self-attention knowledge, in addition to the attention distributions (i.e., the scaled dot-product of queries and keys) that have been used in existing works. Moreover, we show that introducing a teacher assistant (Mirzadeh et al., 2019) also helps the distillation of large pre-trained Transformer models. Experimental results demonstrate that our monolingual model outperforms state-of-the-art baselines in different parameter size of student models. In particular, it retains more than 99% accuracy on SQuAD 2.0 and several GLUE benchmark tasks using 50% of the Transformer parameters and computations of the teacher model. We also obtain competitive results in applying deep self-attention distillation to multilingual pre-trained models.
Citations
Cited by
- TEmBed-T: A Multi-Dimensional Benchmark for Table-Level Embeddings
- Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
- Hierarchical Geometry of Cognitive States in Transformer Embedding Spaces
- Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention
- Interpretable Column Annotation with LLM-Symbolized Decision Process Materialization
- Agentic Graph Retrieval-Augmented Generation for Auditable Commercial Registry Analysis
- Binge Watch: Reproducible Multimodal Benchmarks Datasets for Large-Scale Movie Recommendation on MovieLens-10M and 20M
- Semantic Refinement with LLMs for Graph Representations
- Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
- Graph-based Nearest Neighbors with Dynamic Updates via Random Walks
- From Personalization to Prejudice: Bias and Discrimination in Memory-Enhanced AI Agents for Recruitment
- Chorus: Harmonizing Context and Sensing Signals for Data-Free Model Customization in IoT
- Task Matrices: Linear Maps for Cross-Model Finetuning Transfer
- TiME: Tiny Monolingual Encoders for Efficient NLP Pipelines
- Semantic Grounding Index: Geometric Bounds on Context Engagement in RAG Systems
- Intelligent Scientific Literature Explorer using Machine Learning (ISLE)
- WebOperator: Action-Aware Tree Search for Autonomous Agents in Web Environment
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- Semantic Reconstruction of Adversarial Plagiarism: A Context-Aware Framework for Detecting and Restoring "Tortured Phrases" in Scientific Literature
- CIEGAD: Cluster-Conditioned Interpolative and Extrapolative Framework for Geometry-Aware and Domain-Aligned Data Augmentation
- Interpretation as Linear Transformation: A Cognitive-Geometric Model of Belief and Meaning
- The Erosion of LLM Signatures: Can We Still Distinguish Human and LLM-Generated Scientific Ideas After Iterative Paraphrasing?
- Adaptation of Embedding Models to Financial Filings via LLM Distillation
- Prompting-in-a-Series: Psychology-Informed Contents and Embeddings for Personality Recognition With Decoder-Only Models
- TopiCLEAR: Topic extraction by CLustering Embeddings with Adaptive dimensional Reduction
- Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning
- Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- Flowchart2Mermaid: A Vision-Language Model Powered System for Converting Flowcharts into Editable Diagram Code
- OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
- SAGE: Semantic-Aware Gray-Box Game Regression Testing with Large Language Models
- Breaking It Down: Domain-Aware Semantic Segmentation for Retrieval Augmented Generation
- CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
- LLM-Generated Counterfactual Stress Scenarios for Portfolio Risk Simulation via Hybrid Prompt-RAG Pipeline
- From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures
- BAMAS: Structuring Budget-Aware Multi-Agent Systems
- WhiteningBERT: An Easy Unsupervised Sentence Embedding Approach
- MADRA: Multi-Agent Debate for Risk-Aware Embodied Planning
- Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models
- Building Domain-Specific Small Language Models via Guided Data Generation
- ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization
- When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA
- Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
- The Shifting Landscape of Vaccine Discourse: Insights From a Decade of Pre- to Post-COVID-19 Vaccine Posts on Social Media
- TurkColBERT: A Benchmark of Dense and Late-Interaction Models for Turkish Information Retrieval
- Unifying points of interest taxonomies: mapping OpenStreetMap tags to the Foursquare category system
- How to Select One Among All? An Extensive Empirical Study Towards the Robustness of Knowledge Distillation in Natural Language Understanding
- Distilling Linguistic Context for Language Model Compression
- A Systematic Study of Model Extraction Attacks on Graph Foundation Models
- Black-Box On-Policy Distillation of Large Language Models
- H-Model: Dynamic Neural Architectures for Adaptive Processing
- Think Before You Retrieve: Learning Test-Time Adaptive Search with Small Language Models
- Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
- RAG-targeted Adversarial Attack on LLM-based Threat Detection and Mitigation Framework
- SARCH: Multimodal Search for Archaeological Archives
- A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
- Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification
- CARMA: Comprehensive Automatically-annotated Reddit Mental Health Dataset for Arabic
- Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
- SpecAware: A Spectral-Content Aware Foundation Model for Unifying Multi-Sensor Learning in Hyperspectral Remote Sensing Mapping
- A Memory-Efficient Retrieval Architecture for RAG-Enabled Wearable Medical LLMs-Agents
- Elastic Architecture Search for Efficient Language Models
- Vectorized Context-Aware Embeddings for GAT-Based Collaborative Filtering
- Distilling Multilingual Vision-Language Models: When Smaller Models Stay Multilingual
- NetEcho: From Real-World Streaming Side-Channels to Full LLM Conversation Recovery
- Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
- RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning
- DynaBERT: Dynamic BERT with Adaptive Width and Depth
- Individual differences in neural event segmentation of continuous experiences
- Bi Directional Feedback Fusion for Activity Aware Forecasting of Indoor CO2 and PM2.5
- NERdME: a Named Entity Recognition Dataset for Indexing Research Artifacts in Code Repositories
- EUDAIMONIA: Evaluating Undesirable Dynamics in AI
- Structure of scientific knowledge flows to intergovernmental organizations
- SemOpt: LLM-Driven Code Optimization via Rule-Based Analysis
- DynaStride: Dynamic Stride Windowing with MMCoT for Instructional Multi-Scene Captioning
- Minimizing Human Intervention in Online Classification
- COOPERA: Continual Open-Ended Human-Robot Assistance
- RuSentEval: Linguistic Source, Encoder Force!
- Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings
- The Virtues of Brevity: Avoid Overthinking in Parallel Test-Time Reasoning
- Leveraging semantic similarity for experimentation with AI-generated treatments
- The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts
- Restoring Pruned Large Language Models via Lost Component Compensation
- EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and Retrieval
- Unifying Inductive, Cross-Domain, and Multimodal Learning for Robust and Generalizable Recommendation
- PP3D: An In-Browser Vision-Based Defense Against Web Behavior Manipulation Attacks
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAM
- Robustness in Text-Attributed Graph Learning: Insights, Trade-offs, and New Defenses
- ImaGGen: Zero-Shot Generation of Co-Speech Semantic Gestures Grounded in Language and Image Input
- Accelerating Mobile Language Model via Speculative Decoding and NPU-Coordinated Execution
- Fantastic (small) Retrievers and How to Train Them: mxbai-edge-colbert-v0 Tech Report
- BitNet Distillation
- Simple Projection Variants Improve ColBERT Performance
- CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
- A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
- FinVet: A Collaborative Framework of RAG and External Fact-Checking Agents for Financial Misinformation Detection
- LLM-Oriented Token-Adaptive Knowledge Distillation
- RAG-Pull: Imperceptible Attacks on RAG Systems for Code Generation
- LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding
- Exploring the Promises of Transformer-Based LMs for the Representation of Normative Claims in the Legal Domain
- LinearRAG: Linear Graph Retrieval Augmented Generation on Large-scale Corpora
- From Birdwatch to Community Notes, from Twitter to X: four years of community-based content moderation
- Bridging the Semantic Gap: Contrastive Rewards for Multilingual Text-to-SQL with GRPO
- Personalize Before Retrieve: LLM-based Personalized Query Expansion for User-Centric Retrieval
- FedL2T: Personalized Federated Learning with Two-Teacher Distillation for Seizure Prediction
- ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
- SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAG
- LightMBERT: A Simple Yet Effective Method for Multilingual BERT Distillation
- Multi-Task Pre-Finetuning of Lightweight Transformer Encoders for Text Classification and NER
- Lean Finder: Semantic Search for Mathlib That Understands User Intents
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
- Gamma Mixture Modeling for Cosine Similarity in Small Language Models
- Quantitative Certification of Agentic Tool Selection
- Towards an Efficient, Customizable, and Accessible AI Tutor
- mMARCO: A Multilingual Version of the MS MARCO Passage Ranking Dataset
- Knowledge Graph-Guided Multi-Agent Distillation for Reliable Industrial Question Answering with Datasets
- A Simple but Effective Elaborative Query Reformulation Approach for Natural Language Recommendation
- Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving
- Memory-Augmented Log Analysis with Phi-4-mini: Enhancing Threat Detection in Structured Security Logs
- Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation
- SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligence
- Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- Better with Less: Small Proprietary Models Surpass Large Language Models in Financial Transaction Understanding
- The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
- Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns
- How Well Do LLMs Imitate Human Writing Style?
- PEARL: Peer-Enhanced Adaptive Radio via On-Device LLM
- RestoRect: Degraded Image Restoration via Latent Rectified Flow & Feature Distillation
- The Geometry of Creative Variability: How Credal Sets Expose Calibration Gaps in Language Models
- Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention
- Machine Reading Comprehension: The Role of Contextualized Language Models and Beyond
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- Redefining Machine Simultaneous Interpretation: From Incremental Translation to Human-Like Strategies
- MonoCon: A general framework for learning ultra-compact high-fidelity representations using monotonicity constraints
- Acoustic-based Gender Differentiation in Speech-aware Language Models
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
- PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
- Learning Contextual Retrieval for Robust Conversational Search
- Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
- GGC: Selective Query Correction for Reliable Text-to-SPARQL Generation
- ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs
- Rethinking Network Pruning -- under the Pre-train and Fine-tune Paradigm
- KDLSQ-BERT: A Quantized Bert Combining Knowledge Distillation with Learned Step Size Quantization
- Improving Task-Agnostic BERT Distillation with Layer Mapping Search
- AgentInit: Initializing LLM-based Multi-Agent Systems via Diversity and Expertise Orchestration for Effective and Efficient Collaboration
- Text Meets Topology: Rethinking Out-of-distribution Detection in Text-Rich Networks
- CALL: Context-Aware Low-Latency Retrieval in Disk-Based Vector Databases
- Investigating Traffic Accident Detection Using Multimodal Large Language Models
- Mitigating Strategy-Selection Bias in Reasoning for More Effective Test-Time Scaling
- AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
- InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding
- KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
- Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
- DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models
- AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
- SitLLM: Large Language Models for Sitting Posture Health Understanding via Pressure Sensor Data
- LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
- A Transformer-Based Cross-Platform Analysis of Public Discourse on the 15-Minute City Paradigm
- Difficulty-Aware Agentic Orchestration for Query-Specific Multi-Agent Workflows
- TernaryBERT: Distillation-aware Ultra-low Bit BERT
- JU-NLP at Touché: Covert Advertisement in Conversational AI-Generation and Detection Strategies
- One Teacher is Enough? Pre-trained Language Model Distillation from Multiple Teachers
- Retrieval-Augmented Generation for Reliable Interpretation of Radio Regulations
- Quality Assessment of Tabular Data using Large Language Models and Code Generation
- Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals
- How Small Transformation Expose the Weakness of Semantic Similarity Measures
- When Code Crosses Borders: A Security-Centric Study of LLM-based Code Translation
- mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
- REMI: A Novel Causal Schema Memory Architecture for Personalized Lifestyle Recommendation Agents
- Dynamic Knowledge Distillation for Pre-trained Language Models
- GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation
- BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users
- KGRAG-SC: Knowledge Graph RAG-Assisted Semantic Communication
- IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations
- Towards Open-World Retrieval-Augmented Generation on Knowledge Graph: A Multi-Agent Collaboration Framework
- ART: Adaptive Resampling-based Training for Imbalanced Classification
- Dual-Model Weight Selection and Self-Knowledge Distillation for Medical Image Classification
- SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
- AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark
- Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation
- Granite Embedding R2 Models
- CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering
- LLM-Guided Genetic Improvement: Envisioning Semantic Aware Automated Software Evolution
- An Empirical Study of Knowledge Distillation for Code Understanding Tasks
- Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels
- Synthetic Adaptive Guided Embeddings (SAGE): A Novel Knowledge Distillation Method
- LLMs and Agentic AI in Insurance Decision-Making: Opportunities and Challenges For Africa
- Enhancing Question Generation with Commonsense Knowledge
- Towards Unified Multimodal Financial Forecasting: Integrating Sentiment Embeddings and Market Indicators via Cross-Modal Attention
- SEA-BED: Southeast Asia Embedding Benchmark
- Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems
- Bridging Solidity Evolution Gaps: An LLM-Enhanced Approach for Smart Contract Compilation Error Resolution
- Improving OCR for Historical Texts of Multiple Languages
- Prompt-and-Check: Using Large Language Models to Evaluate Communication Protocol Compliance in Simulation-Based Training
- From Source to Target: Leveraging Transfer Learning for Predictive Process Monitoring in Organizations
- LoSemB: Logic-Guided Semantic Bridging for Inductive Tool Retrieval
- HGMF: A Hierarchical Gaussian Mixture Framework for Scalable Tool Invocation within the Model Context Protocol
- Improving Document Retrieval Coherence for Semantically Equivalent Queries
- CLAP: Coreference-Linked Augmentation for Passage Retrieval
- Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
- Uncovering drivers of climate research in policy with pretrained language models
- Fine-Tuning Small Language Models (SLMs) for Autonomous Web-based Geographical Information Systems (AWebGIS)
- Adaptive Content Restriction for Large Language Models via Suffix Optimization
- End-to-End Personalization: Unifying Recommender Systems with Large Language Models
- Automating AI Failure Tracking: Semantic Association of Reports in AI Incident Database
Related