Learning both Weights and Connections for Efficient Neural Networks
2015/06/08 by Song Han, Han, Song, Jeff Pool +5 · 326 citations
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #cs.CV #cs.LG #cs.NE
paper · pdf · doi:10.48550/arxiv.1506.02626
Published as a conference paper at NIPS 2015
openalex publication_date 2015/06/08 · arxiv created 2015/10/30 · arxiv updated 2015/11/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Neural networks are both computationally intensive and memory intensive, making them difficult to deploy on embedded systems. Also, conventional networks fix the architecture before training starts; as a result, training cannot improve the architecture. To address these limitations, we describe a method to reduce the storage and computation required by neural networks by an order of magnitude without affecting their accuracy by learning only the important connections. Our method prunes redundant connections using a three-step method. First, we train the network to learn which connections are important. Next, we prune the unimportant connections. Finally, we retrain the network to fine tune the weights of the remaining connections. On the ImageNet dataset, our method reduced the number of parameters of AlexNet by a factor of 9x, from 61 million to 6.7 million, without incurring accuracy loss. Similar experiments with VGG-16 found that the number of parameters can be reduced by 13x, from 138 million to 10.3 million, again with no loss of accuracy.
Citations
Cited by
- A Proximal-Gradient Method for Solving Regularized Optimization Problems with General Constraints
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- CosineGate: Semantic Dynamic Routing via Cosine Incompatibility in Residual Networks
- Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?
- Pruning as a Game: Equilibrium-Driven Sparsification of Neural Networks
- Learning to Sense for Driving: Joint Optics-Sensor-Model Co-Design for Semantic Segmentation
- Metanetworks as Regulatory Operators: Learning to Edit for Requirement Compliance
- CoDeQ: End-to-End Joint Model Compression with Dead-Zone Quantizer for High-Sparsity and Low-Precision Networks
- KD-PINN: Knowledge-Distilled PINNs for ultra-low-latency real-time neural PDE solvers
- Evaluating Singular Value Thresholds for DNN Weight Matrices based on Random Matrix Theory
- Effective Fine-Tuning with Eigenvector Centrality Based Pruning
- TPV: Parameter Perturbations Through the Lens of Test Prediction Variance
- SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
- Multi-Granular Node Pruning for Causal Circuit Discovery
- Optimizing LLMs Using Quantization for Mobile Execution
- Neural expressiveness for beyond importance model compression
- HSCP: A Two-Stage Spectral Clustering Framework for Resource-Constrained UAV Identification
- Data-Free Pruning of Self-Attention Layers in LLMs
- PrunedCaps: A Case For Primary Capsules Discrimination
- The brain-AI convergence: Predictive and generative world models for general-purpose computation
- SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot
- Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
- Real-Time Object Tracking with On-Device Deep Learning for Adaptive Beamforming in Dynamic Acoustic Environments
- Post-Pruning Accuracy Recovery via Data-Free Knowledge Distillation
- VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking
- Extreme Model Compression for Edge Vision-Language Models: Sparse Temporal Token Fusion and Adaptive Neural Compression
- A Systematic Study of Compression Ordering for Large Language Models
- Exploiting the Experts: Unauthorized Compression in MoE-LLMs
- E3-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
- Teacher-Guided One-Shot Pruning via Context-Aware Knowledge Distillation
- BD-Net: Has Depth-Wise Convolution Ever Been Applied in Binary Neural Networks?
- EfficientSAM3: Progressive Hierarchical Distillation for Video Concept Segmentation from SAM1, 2, and 3
- Weight Concentration Regularization for Improving Pruning Robustness Under High Sparsity
- FLClear: Visually Verifiable Multi-Client Watermarking for Federated Learning
- Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
- Explore and Establish Synergistic Effects Between Weight Pruning and Coreset Selection in Neural Network Training
- Stratified Knowledge-Density Super-Network for Scalable Vision Transformers
- Beyond One-Way Pruning: Bidirectional Pruning-Regrowth for Extreme Accuracy-Sparsity Tradeoff
- Quantizing Whisper-small: How design choices affect ASR performance
- Pruning as Regularization: Sensitivity-Aware One-Shot Pruning in ASR
- Learning Quantized Continuous Controllers for Integer Hardware
- MI-to-Mid Distilled Compression (M2M-DC): An Hybrid-Information-Guided-Block Pruning with Progressive Inner Slicing Approach to Model Compression
- ML-EcoLyzer: Quantifying the Environmental Cost of Machine Learning Inference Across Frameworks and Hardware
- CAMP-HiVe: Cyclic Pair Merging based Efficient DNN Pruning with Hessian-Vector Approximation for Resource-Constrained Systems
- An Efficient Gradient-Aware Error-Bounded Lossy Compressor for Federated Learning
- Efficient CNN Inference on Ultra-Low-Power MCUs via Saturation-Aware Convolution
- FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
- TwIST: Rigging the Lottery in Transformers with Independent Subnetwork Training
- SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking
- FPS: Feedforward-based Parameter Selection For Efficient Fine-Tuning
- Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models
- Safe Screening Rules for Group SLOPE
- Emergent Sparsity in Frozen Random CNN Feature Extractors for Deep Reinforcement Learning
- Lipschitz-aware Linearity Grafting for Certified Robustness
- TsetlinKWS: A 65nm 16.58uW, 0.63mm2 State-Driven Convolutional Tsetlin Machine-Based Accelerator For Keyword Spotting
- Adaptive Training of INRs via Pruning and Densification
- Efficient Cost-and-Quality Controllable Arbitrary-scale Super-resolution with Fourier Constraints
- TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
- Frustratingly Easy Task-aware Pruning for Large Language Models
- GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
- A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
- KARIPAP: Quantum-Inspired Tensor Network Compression of Large Language Models Using Infinite Projected Entangled Pair States and Tensor Renormalization Group
- VeFA: Vector-Based Feature Space Adaptation for Robust Model Fine-Tuning
- C-SWAP: Explainability-Aware Structured Pruning for Efficient Neural Networks Compression
- S2AP: Score-space Sharpness Minimization for Adversarial Pruning
- Elastic ViTs from Pretrained Models without Retraining
- Self-Evidencing Through Hierarchical Gradient Decomposition: A Dissipative System That Maintains Non-Equilibrium Steady-State by Minimizing Variational Free Energy
- ALPINE: Closed-Loop Adaptive Privacy Budget Allocation for Mobile Edge Crowdsensing
- A Multimodal Approach to Heritage Preservation in the Context of Climate Change
- Convergence, design and training of continuous-time dropout as a random batch method
- Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
- Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models
- Structured Sparsity and Weight-adaptive Pruning for Memory and Compute efficient Whisper models
- Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
- Pruning Cannot Hurt Robustness: Certified Trade-offs in Reinforcement Learning
- Optimally Deep Networks -- Adapting Model Depth to Datasets for Superior Efficiency
- CauchyNet: Compact and Data-Efficient Learning using Holomorphic Activation Functions
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
- Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
- Shift-Invariant Attribute Scoring for Kolmogorov-Arnold Networks via Shapley Value
- Fewer Weights, More Problems: A Practical Attack on LLM Pruning
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot
- Downsized and Compromised?: Assessing the Faithfulness of Model Compression
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Expand Neurons, Not Parameters
- PolyKAN: A Polyhedral Analysis Framework for Provable and Approximately Optimal KAN Compression
- OptiFLIDS: Optimized Federated Learning for Energy-Efficient Intrusion Detection in IoT
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
- CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
- Interpret, prune and distill Donut : towards lightweight VLMs for VQA on document
- Differentiable Sparsity via D-Gating: Simple and Versatile Structured Penalization
- PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
- Lightweight error mitigation strategies for post-training N:M activation sparsity in LLMs
- Deep learning for interval-censored failure time data from case-cohort studies
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
- Budgeted Broadcast: An Activity-Dependent Pruning Rule for Neural Network Efficiency
- StructPrune: Structured Global Pruning asymptotics with O(√(N)) GPU Memory
- Fine-tuning of Large Language Models for Domain-Specific Cybersecurity Knowledge
- RAM-NAS: Resource-aware Multiobjective Neural Architecture Search Method for Robot Vision Tasks
- Smaller is Better: Enhancing Transparency in Vehicle AI Systems via Pruning
- Recursive transformers for semiconductor thermo-mechanical reliability
- Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
- GradMAP: Faster Layer Pruning with Gradient Metric and Projection Compensation
- Shared-Weights Extender and Gradient Voting for Neural Network Expansion
- Do Sparse Subnetworks Exhibit Cognitively Aligned Attention? Effects of Pruning on Saliency Map Fidelity, Sparsity, and Concept Coherence
- Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models
- Convolutional Neural Network Optimization for Beehive Classification Using Bioacoustic Signals
- Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
- Closing the Loop Inside Neural Networks: Causality-Guided Layer Adaptation for Fault Recovery Control
- DISPATCH: Distilling Selective Patches for Speech Enhancement
- TinySR: Pruning Diffusion for Real-World Image Super-Resolution
- FedERL: Federated Efficient and Robust Learning for Common Corruptions
- The Lifecycle Principle: Stabilizing Dynamic Neural Networks with State Memory
- SugarcaneShuffleNet: A Very Fast, Lightweight Convolutional Neural Network for Diagnosis of 15 Sugarcane Leaf Diseases
- Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
- GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
- Boosted Training of Lightweight Early Exits for Optimizing CNN Image Classification Inference
- HufuNet: Embedding the Left Piece as Watermark and Keeping the Right Piece for Ownership Verification in Deep Neural Networks
- Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruning
- Electricity Demand and Grid Impacts of AI Data Centers: Challenges and Prospects
- Explaining How Quantization Disparately Skews a Model
- Delta Activations: A Representation for Finetuned Large Language Models
- Integrating Pruning with Quantization for Efficient Deep Neural Networks Compression
- Learning Neural Decoding with Parallelism and Self-Coordination for Quantum Error Correction
- FastCaps: A Design Methodology for Accelerating Capsule Network on Field Programmable Gate Arrays
- Silent Until Sparse: Backdoor Attacks on Semi-Structured Sparsity
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- Application of discrete Ricci curvature in pruning randomly wired neural networks: A case study with chest x-ray classification of COVID-19
- Deep Learning for Personalized Binaural Audio Reproduction
- Theory Foundation of Physics-Enhanced Residual Learning
- AI Compute Architecture and Evolution Trends
- Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
- Low-bit Model Quantization for Deep Neural Networks: A Survey
- FedUP: Efficient Pruning-based Federated Unlearning for Model Poisoning Attacks
- One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
- Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
- ONG: One-Shot NMF-based Gradient Masking for Efficient Model Sparsification
- Exploring Efficiency Frontiers of Thinking Budget in Medical Reasoning: Scaling Laws between Computational Resources and Reasoning Quality
- FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
- EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models
- Synaptic Pruning: A Biological Inspiration for Deep Learning Regularization
- Harnessing Input-Adaptive Inference for Efficient VLN
- Neural Tangent Knowledge Distillation for Optical Convolutional Networks
- Investigating 1-Bit Quantization in Transformer-Based Top Tagging
- Why Does Stochastic Gradient Descent Slow Down in Low-Precision Training?
- Crisp Attention: Regularizing Transformers via Structured Sparsity
- Efficient Deep Neural Receiver with Post-Training Quantization
- Optimal Brain Connection: Towards Efficient Structural Pruning
- Pruning Large Language Models by Identifying and Preserving Functional Networks
- Advanced Hybrid Transformer LSTM Technique with Attention and TS Mixer for Drilling Rate of Penetration Prediction
- TopKD: Top-scaled Knowledge Distillation
- CARD: A Cache-Assisted Parallel Speculative Decoding Framework via Query-and-Correct Paradigm for Accelerating LLM Inference
- Anticipating Decoherence: a Predictive Framework for Enhancing Coherence in Quantum Emitters
- FAIR-Pruner: Leveraging Tolerance of Difference for Flexible Automatic Layer-Wise Neural Network Pruning
- FedAPTA: Federated Multi-task Learning for Heterogeneous Devices with Adaptive Layer-wise Pruning and Task-aware Aggregation
- Uncertainty Quantification for Large-Scale Deep Networks via Post-StoNet Modeling
- On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
- Large-Scale Evolution of Image Classifiers
- Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers
- LCS: An AI-based Low-Complexity Scaler for Power-Efficient Super-Resolution of Game Content
- FGFP: A Fractional Gaussian Filter and Pruning for Deep Neural Networks Compression
- Compression Strategies for Efficient Multimodal LLMs in Medical Contexts
- LinDeps: A Fine-tuning Free Post-Pruning Method to Remove Layer-Wise Linear Dependencies with Guaranteed Performance Preservation
- Hot-Swap MarkBoard: An Efficient Black-box Watermarking Approach for Large-scale Model Distribution
- Leveraging Sparse Linear Layers for Debuggable Deep Networks
- Knowledge Grafting: A Mechanism for Optimizing AI Model Deployment in Resource-Constrained Environments
- CD-SGD: Distributed Stochastic Gradient Descent with Compression and Delay Compensation
- FINN-R: An End-to-End Deep-Learning Framework for Fast Exploration of Quantized Neural Networks
- The Benchmark Illusion: Pruned LLMs Can Pass Multiple Choice but Fail to Answer
- A GPU-Outperforming FPGA Accelerator Architecture for Binary Convolutional Neural Networks
- EPNAS: Efficient Progressive Neural Architecture Search
- Online Training and Pruning of Deep Reinforcement Learning Networks
- COLI: A Hierarchical Efficient Compressor for Large Images
- On the Stability of the Jacobian Matrix in Deep Neural Networks
- Towards Design Methodology of Efficient Fast Algorithms for Accelerating Generative Adversarial Networks on FPGAs
- Olica: Efficient Structured Pruning of Large Language Models without Retraining
- Reliable Identification of Redundant Kernels for Convolutional Neural Network Compression
- SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
- FLoRIST: Singular Value Thresholding for Efficient and Accurate Federated Fine-Tuning of Large Language Models
- Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
- Text Embedding Knows How to Quantize Text-Guided Diffusion Models
- WGLE:Backdoor-free and Multi-bit Black-box Watermarking for Graph Neural Networks
- Spatial Lifting for Dense Prediction
- Effects of relational graph modularity and depth on the learning performance of neural networks
- Detecting and Pruning Prominent but Detrimental Neurons in Large Language Models
- Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R
- Interpretability-Aware Pruning for Efficient Medical Image Analysis
- Lightweight Object Detection Using Quantized YOLOv4-Tiny for Emergency Response in Aerial Imagery
- UnIT: Scalable Unstructured Inference-Time Pruning for MAC-efficient Neural Inference on MCUs
- Weight-Space Mixture-of-Experts for Implicit Neural Representation Classification
- AdaDPIGU: Differentially Private SGD with Adaptive Clipping and Importance-Based Gradient Updates for Deep Neural Networks
- Model Compression using Progressive Channel Pruning
- DANCE: Resource-Efficient Neural Architecture Search with Data-Aware and Continuous Adaptation
- TinyProto: Communication-Efficient Federated Learning with Sparse Prototypes in Resource-Constrained Environments
- Scalable Interconnect Learning in Boolean Networks
- High-Layer Attention Pruning with Rescaling
- Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization
- evMLP: An Efficient Event-Driven MLP Architecture for Vision
- LotteryCodec: Searching the Implicit Representation in a Random Network for Low-Complexity Image Compression
- Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
- CS-VLM: Compressed Sensing Attention for Efficient Vision-Language Representation Learning
- Hyperpruning: Efficient Search through Pruned Variants of Recurrent Neural Networks Leveraging Lyapunov Spectrum
- Phase-Only Positioning: Overcoming Integer Ambiguity Challenge through Deep Learning
- SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark
- Towards Universal & Efficient Model Compression via Exponential Torque Pruning
- Projected Compression: Trainable Projection for Efficient Transformer Compression
- PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
- Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
- Loss-Aware Automatic Selection of Structured Pruning Criteria for Deep Neural Network Acceleration
- DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
- Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using Sparsity
- Revisiting LoRA through the Lens of Parameter Redundancy: Spectral Encoding Helps
- Fast and Accurate, Convolutional Neural Network Based Approach for Object Detection from UAV
- SparseDPD: A Sparse Neural Network-based Digital Predistortion FPGA Accelerator for RF Power Amplifier Linearization
- NetSenseML: Network-Adaptive Compression for Efficient Distributed Machine Learning
- Efficient and Privacy-Preserving Soft Prompt Transfer for LLMs
- SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity
- Understanding LLM-Centric Challenges for Deep Learning Frameworks: An Empirical Analysis
- Dynamic Acoustic Model Architecture Optimization in Training for ASR
- Studying the Consistency and Composability of Lottery Ticket Pruning Masks
- Gradient-based Fine-Tuning through Pre-trained Model Regularization
- Compression Aware Certified Training
- Machine Unlearning for Robust DNNs: Attribution-Guided Partitioning and Neuron Pruning in Noisy Environments
- Dynamic Sparse Training of Diagonally Sparse Networks
- SecONNds: Secure Outsourced Neural Network Inference on ImageNet
- A Layered Self-Supervised Knowledge Distillation Framework for Efficient Multimodal Learning on the Edge
- SAFE: Finding Sparse and Flat Minima to Improve Pruning
- Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation
- Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias
- Event Classification of Accelerometer Data for Industrial Package Monitoring with Embedded Deep Learning
- Structured Pruning and Quantization for Learned Image Compression
- The Promise of Spiking Neural Networks for Ubiquitous Computing: A Survey and New Perspectives
- Exchangeability in Neural Network and its Application to Dynamic Pruning
- Memory-Efficient FastText: A Comprehensive Approach Using Double-Array Trie Structures and Mark-Compact Memory Management
- Knowledge Distillation: A Survey
- It Takes a Good Model to Train a Good Model: Generalized Gaussian Priors for Optimized LLMs
- A Simple Linear Patch Revives Layer-Pruned Large Language Models
- Smooth Model Compression without Fine-Tuning
- LPASS: Linear Probes as Stepping Stones for vulnerability detection using compressed LLMs
- Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1
- TSENOR: Highly-Efficient Algorithm for Finding Transposable N:M Sparse Masks
- Learning Interpretable Differentiable Logic Networks for Tabular Regression
- ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
- Interpretable Scaling Behavior in Sparse Subnetwork Representations of Quantum States
- ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
- Node pruning reveals compact and optimal substructures within large networks
- Progressive Data Dropout: An Embarrassingly Simple Approach to Faster Training
- EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices
- Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques
- M-Wanda: Improving One-Shot Pruning for Multilingual LLMs
- LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions
- Global Minimizers of ℓp-Regularized Objectives Yield the Sparsest ReLU Neural Networks
- Efficient Large Language Model Inference with Neural Block Linearization
- Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
- End-to-end fully-binarized network design: from Generic Learned Thermometer to Block Pruning
- WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
- Efficient SRAM-PIM Co-design by Joint Exploration of Value-Level and Bit-Level Sparsity
- PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
- Learning without Isolation: Pathway Protection for Continual Learning
- NeuroTrails: Training with Dynamic Sparse Heads as the Key to Effective Ensembling
- Two-Stage Regularization-Based Structured Pruning for LLMs
- CIM-NET: A Video Denoising Deep Neural Network Model Optimized for Computing-in-Memory Architectures
- TRIM: Achieving Extreme Sparsity with Targeted Row-wise Iterative Metric-driven Pruning
- Cooperative NOMA Meets Emerging Technologies: A Survey for Next-Generation Wireless Networks
- DeFTX: Denoised Sparse Fine-Tuning for Zero-Shot Cross-Lingual Transfer
- Void in Language Models
- Improved Methods for Model Pruning and Knowledge Distillation
- An Empirical Study of the Impact of Hyperparameter Tuning and Model Optimization on the Performance Properties of Deep Neural Networks
- Adaptive Pruning of Deep Neural Networks for Resource-Aware Embedded Intrusion Detection on the Edge
- Efficient Privacy-Preserving Cross-Silo Federated Learning with Multi-Key Homomorphic Encryption
- Optimal Client Sampling in Federated Learning with Client-Level Heterogeneous Differential Privacy
- Energy-Aware Deep Learning on Resource-Constrained Hardware
- Build a Compact Binary Neural Network through Bit-level Sensitivity and Data Pruning
- QVGen: Pushing the Limit of Quantized Video Generative Models
- Adversarially Robust Spiking Neural Networks with Sparse Connectivity
- Generated Images Are Easier to Forget: A Machine Unlearning Perspective for Synthetic Image Detection
- VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits
- BINGO: A Novel Pruning Mechanism to Reduce the Size of Neural Networks
- ILIF: Temporal Inhibitory Leaky Integrate-and-Fire Neuron for Overactivation in Spiking Neural Networks
- PDE: Gene Effect Inspired Parameter Dynamic Evolution for Low-light Image Enhancement
- Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained Environments
- Guiding Evolutionary AutoEncoder Training with Activation-Based Pruning Operators
- Sparse Training from Random Initialization: Aligning Lottery Ticket Masks using Weight Symmetry
- Efficient Shapley Value-based Non-Uniform Pruning of Large Language Models
- Energy-Efficient Neuromorphic Computing for Edge AI: A Framework with Adaptive Spiking Neural Networks and Hardware-Aware Optimization
- Model Connectomes: A Generational Approach to Data-Efficient Language Models
- Sparse Subspace-to-Expert Sharing for Task-Agnostic Continual Learning
- The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning
- Routing the Lottery: Adaptive Subnetworks for Heterogeneous Data
- Personalized Artificial General Intelligence (AGI) via Neuroscience-Inspired Continuous Learning Systems
- 3DPyranet Features Fusion for Spatio-temporal Feature Learning
- Probabilistic fine-tuning of pruning masks and PAC-Bayes self-bounded learning
- Channel-wise Dynamic Knowledge Distillation via Adaptive Sample Generation for Action Recognition
- Event-Based Eye Tracking. 2025 Event-based Vision Workshop
- The Rise of Small Language Models in Healthcare: A Comprehensive Survey
- Rethinking Reservoir Pruning: A Dynamical Perspective for Echo State Networks
- StaticSegFormer: An Efficient High-Performance Semantic Segmentation Based on Static Structured Pruning
- LaPrune: Controllable Differentiable Sparsity at Million Scale
- Understanding Fault Tolerance of Adversarially Robust Pruned Models
- BackSlash: Rate Constrained Optimized Training of Large Language Models
- The Neural Pruning Law Hypothesis
- Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
- TrojanDam: Detection-Free Backdoor Defense in Federated Learning through Proactive Model Robustification utilizing OOD Data
- Efficient Adaptation of Deep Neural Networks for Semantic Segmentation in Space Applications
- Connecting Parameter Magnitudes and Hessian Eigenspaces at Scale using Sketched Methods
- NoWag: A Unified Framework for Shape Preserving Compression of Large Language Models
- Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
- Collaborative Learning of On-Device Small Model and Cloud-Based Large Model: Advances and Future Directions
- You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models
- APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning
- Effective pruning of task-trained recurrent neural networks using noisy fluctuations and connection rescaling
- High-Efficiency Split Computing for Cooperative Edge Systems: A Novel Compressed Sensing Bottleneck
- Towards Unbiased Federated Graph Learning: Label and Topology Perspectives
- Learning with Spike Synchrony in Spiking Neural Networks
- Early-Bird Diffusion: Investigating and Leveraging Timestep-Aware Early-Bird Tickets in Diffusion Models for Efficient Training
- A Multicore and Edge TPU-Accelerated Multimodal TinyML System for Livestock Behavior Recognition
- Generative Artificial Intelligence for Internet of Things Computing: A Systematic Survey
- Heuristic Methods are Good Teachers to Distill MLPs for Graph Link Prediction
- Two is Better than One: Efficient Ensemble Defense for Robust and Compact Models
Related