Learning both Weights and Connections for Efficient Neural Networks
2015/06/08 by Song Han, Han, Song, Jeff Pool +5 · 170 citations
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)
paper · pdf · doi:10.48550/arxiv.1506.02626
openalex publication_date 2015/06/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Neural networks are both computationally intensive and memory intensive, making them difficult to deploy on embedded systems. Also, conventional networks fix the architecture before training starts; as a result, training cannot improve the architecture. To address these limitations, we describe a method to reduce the storage and computation required by neural networks by an order of magnitude without affecting their accuracy by learning only the important connections. Our method prunes redundant connections using a three-step method. First, we train the network to learn which connections are important. Next, we prune the unimportant connections. Finally, we retrain the network to fine tune the weights of the remaining connections. On the ImageNet dataset, our method reduced the number of parameters of AlexNet by a factor of 9x, from 61 million to 6.7 million, without incurring accuracy loss. Similar experiments with VGG-16 found that the number of parameters can be reduced by 13x, from 138 million to 10.3 million, again with no loss of accuracy.
Citations
Cited by
- A Proximal-Gradient Method for Solving Regularized Optimization Problems with General Constraints
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- CosineGate: Semantic Dynamic Routing via Cosine Incompatibility in Residual Networks
- Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?
- Pruning as a Game: Equilibrium-Driven Sparsification of Neural Networks
- Learning to Sense for Driving: Joint Optics-Sensor-Model Co-Design for Semantic Segmentation
- Metanetworks as Regulatory Operators: Learning to Edit for Requirement Compliance
- CoDeQ: End-to-End Joint Model Compression with Dead-Zone Quantizer for High-Sparsity and Low-Precision Networks
- KD-PINN: Knowledge-Distilled PINNs for ultra-low-latency real-time neural PDE solvers
- Evaluating Singular Value Thresholds for DNN Weight Matrices based on Random Matrix Theory
- Effective Fine-Tuning with Eigenvector Centrality Based Pruning
- TPV: Parameter Perturbations Through the Lens of Test Prediction Variance
- SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
- Multi-Granular Node Pruning for Causal Circuit Discovery
- Optimizing LLMs Using Quantization for Mobile Execution
- Neural expressiveness for beyond importance model compression
- HSCP: A Two-Stage Spectral Clustering Framework for Resource-Constrained UAV Identification
- Data-Free Pruning of Self-Attention Layers in LLMs
- PrunedCaps: A Case For Primary Capsules Discrimination
- The brain-AI convergence: Predictive and generative world models for general-purpose computation
- Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
- Real-Time Object Tracking with On-Device Deep Learning for Adaptive Beamforming in Dynamic Acoustic Environments
- Post-Pruning Accuracy Recovery via Data-Free Knowledge Distillation
- VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking
- Extreme Model Compression for Edge Vision-Language Models: Sparse Temporal Token Fusion and Adaptive Neural Compression
- A Systematic Study of Compression Ordering for Large Language Models
- Exploiting the Experts: Unauthorized Compression in MoE-LLMs
- E3-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
- Teacher-Guided One-Shot Pruning via Context-Aware Knowledge Distillation
- BD-Net: Has Depth-Wise Convolution Ever Been Applied in Binary Neural Networks?
- EfficientSAM3: Progressive Hierarchical Distillation for Video Concept Segmentation from SAM1, 2, and 3
- Weight Variance Amplifier Improves Accuracy in High-Sparsity One-Shot Pruning
- FLClear: Visually Verifiable Multi-Client Watermarking for Federated Learning
- Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
- Explore and Establish Synergistic Effects Between Weight Pruning and Coreset Selection in Neural Network Training
- Stratified Knowledge-Density Super-Network for Scalable Vision Transformers
- Beyond One-Way Pruning: Bidirectional Pruning-Regrowth for Extreme Accuracy-Sparsity Tradeoff
- Quantizing Whisper-small: How design choices affect ASR performance
- Pruning as Regularization: Sensitivity-Aware One-Shot Pruning in ASR
- Learning Quantized Continuous Controllers for Integer Hardware
- MI-to-Mid Distilled Compression (M2M-DC): An Hybrid-Information-Guided-Block Pruning with Progressive Inner Slicing Approach to Model Compression
- ML-EcoLyzer: Quantifying the Environmental Cost of Machine Learning Inference Across Frameworks and Hardware
- CAMP-HiVe: Cyclic Pair Merging based Efficient DNN Pruning with Hessian-Vector Approximation for Resource-Constrained Systems
- An Efficient Gradient-Aware Error-Bounded Lossy Compressor for Federated Learning
- Efficient CNN Inference on Ultra-Low-Power MCUs via Saturation-Aware Convolution
- FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
- TwIST: Rigging the Lottery in Transformers with Independent Subnetwork Training
- SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking
- FPS: Feedforward-based Parameter Selection For Efficient Fine-Tuning
- Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models
- Emergent Sparsity in Frozen Random CNN Feature Extractors for Deep Reinforcement Learning
- Lipschitz-aware Linearity Grafting for Certified Robustness
- TsetlinKWS: A 65nm 16.58uW, 0.63mm2 State-Driven Convolutional Tsetlin Machine-Based Accelerator For Keyword Spotting
- Adaptive Training of INRs via Pruning and Densification
- Efficient Cost-and-Quality Controllable Arbitrary-scale Super-resolution with Fourier Constraints
- TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
- Frustratingly Easy Task-aware Pruning for Large Language Models
- GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
- A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
- KARIPAP: Quantum-Inspired Tensor Network Compression of Large Language Models Using Infinite Projected Entangled Pair States and Tensor Renormalization Group
- VeFA: Vector-Based Feature Space Adaptation for Robust Model Fine-Tuning
- C-SWAP: Explainability-Aware Structured Pruning for Efficient Neural Networks Compression
- S2AP: Score-space Sharpness Minimization for Adversarial Pruning
- Elastic ViTs from Pretrained Models without Retraining
- Self-Evidencing Through Hierarchical Gradient Decomposition: A Dissipative System That Maintains Non-Equilibrium Steady-State by Minimizing Variational Free Energy
- ALPINE: Closed-Loop Adaptive Privacy Budget Allocation for Mobile Edge Crowdsensing
- A Multimodal Approach to Heritage Preservation in the Context of Climate Change
- Convergence, design and training of continuous-time dropout as a random batch method
- Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
- Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models
- Structured Sparsity and Weight-adaptive Pruning for Memory and Compute efficient Whisper models
- Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
- Pruning Cannot Hurt Robustness: Certified Trade-offs in Reinforcement Learning
- Optimally Deep Networks -- Adapting Model Depth to Datasets for Superior Efficiency
- CauchyNet: Compact and Data-Efficient Learning using Holomorphic Activation Functions
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
- Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
- Shift-Invariant Attribute Scoring for Kolmogorov-Arnold Networks via Shapley Value
- Fewer Weights, More Problems: A Practical Attack on LLM Pruning
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot
- Downsized and Compromised?: Assessing the Faithfulness of Model Compression
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Expand Neurons, Not Parameters
- PolyKAN: A Polyhedral Analysis Framework for Provable and Approximately Optimal KAN Compression
- OptiFLIDS: Optimized Federated Learning for Energy-Efficient Intrusion Detection in IoT
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
- CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
- Interpret, prune and distill Donut : towards lightweight VLMs for VQA on document
- Differentiable Sparsity via D-Gating: Simple and Versatile Structured Penalization
- PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
- Lightweight error mitigation strategies for post-training N:M activation sparsity in LLMs
- Deep learning for interval-censored failure time data from case-cohort studies
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
- Budgeted Broadcast: An Activity-Dependent Pruning Rule for Neural Network Efficiency
- StructPrune: Structured Global Pruning asymptotics with O(√(N)) GPU Memory
- Fine-tuning of Large Language Models for Domain-Specific Cybersecurity Knowledge
- RAM-NAS: Resource-aware Multiobjective Neural Architecture Search Method for Robot Vision Tasks
- Smaller is Better: Enhancing Transparency in Vehicle AI Systems via Pruning
- Recursive transformers for semiconductor thermo-mechanical reliability
- Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
- GradMAP: Faster Layer Pruning with Gradient Metric and Projection Compensation
- Shared-Weights Extender and Gradient Voting for Neural Network Expansion
- Do Sparse Subnetworks Exhibit Cognitively Aligned Attention? Effects of Pruning on Saliency Map Fidelity, Sparsity, and Concept Coherence
- Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models
- Convolutional Neural Network Optimization for Beehive Classification Using Bioacoustic Signals
- Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
- Closing the Loop Inside Neural Networks: Causality-Guided Layer Adaptation for Fault Recovery Control
- DISPATCH: Distilling Selective Patches for Speech Enhancement
- TinySR: Pruning Diffusion for Real-World Image Super-Resolution
- FedERL: Federated Efficient and Robust Learning for Common Corruptions
- The Lifecycle Principle: Stabilizing Dynamic Neural Networks with State Memory
- SugarcaneShuffleNet: A Very Fast, Lightweight Convolutional Neural Network for Diagnosis of 15 Sugarcane Leaf Diseases
- Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
- GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
- Boosted Training of Lightweight Early Exits for Optimizing CNN Image Classification Inference
- HufuNet: Embedding the Left Piece as Watermark and Keeping the Right Piece for Ownership Verification in Deep Neural Networks
- Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruning
- Electricity Demand and Grid Impacts of AI Data Centers: Challenges and Prospects
- Explaining How Quantization Disparately Skews a Model
- Delta Activations: A Representation for Finetuned Large Language Models
- Integrating Pruning with Quantization for Efficient Deep Neural Networks Compression
- Learning Neural Decoding with Parallelism and Self-Coordination for Quantum Error Correction
- FastCaps: A Design Methodology for Accelerating Capsule Network on Field Programmable Gate Arrays
- Silent Until Sparse: Backdoor Attacks on Semi-Structured Sparsity
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- Application of discrete Ricci curvature in pruning randomly wired neural networks: A case study with chest x-ray classification of COVID-19
- Deep Learning for Personalized Binaural Audio Reproduction
- Theory Foundation of Physics-Enhanced Residual Learning
- AI Compute Architecture and Evolution Trends
- Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
- FedUP: Efficient Pruning-based Federated Unlearning for Model Poisoning Attacks
- One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
- Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
- ONG: One-Shot NMF-based Gradient Masking for Efficient Model Sparsification
- Exploring Efficiency Frontiers of Thinking Budget in Medical Reasoning: Scaling Laws between Computational Resources and Reasoning Quality
- FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
- EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models
- Synaptic Pruning: A Biological Inspiration for Deep Learning Regularization
- Harnessing Input-Adaptive Inference for Efficient VLN
- Neural Tangent Knowledge Distillation for Optical Convolutional Networks
- Investigating 1-Bit Quantization in Transformer-Based Top Tagging
- Why Does Stochastic Gradient Descent Slow Down in Low-Precision Training?
- Crisp Attention: Regularizing Transformers via Structured Sparsity
- Efficient Deep Neural Receiver with Post-Training Quantization
- Optimal Brain Connection: Towards Efficient Structural Pruning
- Pruning Large Language Models by Identifying and Preserving Functional Networks
- Advanced Hybrid Transformer LSTM Technique with Attention and TS Mixer for Drilling Rate of Penetration Prediction
- TopKD: Top-scaled Knowledge Distillation
- CARD: A Cache-Assisted Parallel Speculative Decoding Framework via Query-and-Correct Paradigm for Accelerating LLM Inference
- Anticipating Decoherence: a Predictive Framework for Enhancing Coherence in Quantum Emitters
- FAIR-Pruner: Leveraging Tolerance of Difference for Flexible Automatic Layer-Wise Neural Network Pruning
- FedAPTA: Federated Multi-task Learning for Heterogeneous Devices with Adaptive Layer-wise Pruning and Task-aware Aggregation
- Uncertainty Quantification for Large-Scale Deep Networks via Post-StoNet Modeling
- On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
- Large-Scale Evolution of Image Classifiers
- Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers
- LCS: An AI-based Low-Complexity Scaler for Power-Efficient Super-Resolution of Game Content
- FGFP: A Fractional Gaussian Filter and Pruning for Deep Neural Networks Compression
- Compression Strategies for Efficient Multimodal LLMs in Medical Contexts
- LinDeps: A Fine-tuning Free Post-Pruning Method to Remove Layer-Wise Linear Dependencies with Guaranteed Performance Preservation
- Hot-Swap MarkBoard: An Efficient Black-box Watermarking Approach for Large-scale Model Distribution
- Leveraging Sparse Linear Layers for Debuggable Deep Networks
- Knowledge Grafting: A Mechanism for Optimizing AI Model Deployment in Resource-Constrained Environments
- CD-SGD: Distributed Stochastic Gradient Descent with Compression and Delay Compensation
Related