SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
2023/01/02 by Elias Frantar, Dan Alistarh, Frantar, Elias +1 · 3 voices · 170 citations
Computer Science · Engineering · #Advanced Data Storage Technologies #Ferroelectric and Negative Capacitance Devices #Topic Modeling #cs.LG
paper · pdf · doi:10.48550/arxiv.2301.00774
openalex publication_date 2023/01/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
We show for the first time that large-scale generative pretrained transformer (GPT) family models can be pruned to at least 50% sparsity in one-shot, without any retraining, at minimal loss of accuracy. This is achieved via a new pruning method called SparseGPT, specifically designed to work efficiently and accurately on massive GPT-family models. We can execute SparseGPT on the largest available open-source models, OPT-175B and BLOOM-176B, in under 4.5 hours, and can reach 60% unstructured sparsity with negligible increase in perplexity: remarkably, more than 100 billion weights from these models can be ignored at inference time. SparseGPT generalizes to semi-structured (2:4 and 4:8) patterns, and is compatible with weight quantization approaches. The code is available at: https://github.com/IST-DASLab/sparsegpt.
Cited by
- Unified Static-Dynamic Pruning for Efficient LLM Inference
- Importance-Aware OBS Pruning for Diffusion Models
- SHUFFLESPARSE: Learned Shuffles for Structured Sparse Networks
- GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding
- CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning
- Generalized Fisher-Weighted SVD: Scalable Kronecker-Factored Fisher Approximation for Compressing Large Language Models
- OrderMoE: An expert similarity driven distributed edge MoE inference
- It Takes a MAESTRO To Prune Bad Experts
- NIRVANA: Structured Pruning Reimagined for Large Language Model Compression
- Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models
- Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI
- Your Language Model Secretly Contains Personality Subnetworks
- Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression
- The 4/δ Bound: Designing Predictable LLM-Verifier Systems for Formal Method Guarantee
- 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
- Position: It's Time to Act on the Risk of Efficient Personalized Text Generation
- A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
- Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
- Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?
- TriSP: Tri-Signal Structured Pruning for Large Language Models
- LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
- TimeBill: Time-Budgeted Inference for Large Language Models
- Pruning as a Game: Equilibrium-Driven Sparsification of Neural Networks
- Making Large Language Models Efficient Dense Retrievers
- Can abstract concepts from LLM improve SLM performance?
- MoE Pathfinder: Trajectory-driven Expert Pruning
- VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
- SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block Skipping
- Resting Neurons, Active Insights: Improving Input Sparsification for Large Language Models
- SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
- ZK-APEX: Zero-Knowledge Approximate Personalized Unlearning with Executable Proofs
- Leveraging KV Similarity for Online Structured Pruning in LLMs
- SQ-format: A Unified Sparse-Quantized Hardware-friendly Data Format for LLMs
- Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates
- Data-Free Pruning of Self-Attention Layers in LLMs
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
- Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
- HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
- EfficientXpert: Efficient Domain Adaptation for Large Language Models via Propagation-Aware Pruning
- Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
- Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- INTERLACE: Interleaved Layer Pruning and Efficient Adaptation in Large Vision-Language Models
- ModHiFi: Identifying High Fidelity predictive components for Model Modification
- FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
- Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models
- Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
- Exploiting the Experts: Unauthorized Compression in MoE-LLMs
- E3-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
- Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers
- Automatic Pruning Discovery for Large Language Models
- PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
- TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZone
- Weight-sparse transformers have interpretable circuits
- MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
- Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
- KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
- A3: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures
- EcoSpa: Efficient Transformer Training with Coupled Sparsity
- APP: Accelerated Path Patching with Task-Specific Pruning
- TwIST: Rigging the Lottery in Transformers with Independent Subnetwork Training
- Optimal Singular Damage: Efficient LLM Inference in Low Storage Regimes
- IG-Pruning: Input-Guided Block Pruning for Large Language Models
- Continual Learning, Not Training: Online Adaptation For Agents
- AI Progress Should Be Measured by Capability-Per-Resource, Not Scale Alone: A Framework for Gradient-Guided Resource Allocation in LLMs
- 1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Models
- NeuronMM: High-Performance Matrix Multiplication for LLM Inference on AWS Trainium
- FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference
- OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
- The Curious Case of In-Training Compression of State Space Models
- Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for On-Device LLMs
- PRO: Enabling Precise and Robust Text Watermark for Open-Source LLMs
- PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization
- TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
- FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
- Frustratingly Easy Task-aware Pruning for Large Language Models
- The Structural Scalpel: Automated Contiguous Layer Pruning for Large Language Models
- Scaling Up Efficient Small Language Models Serving and Deployment for Semantic Job Search
- Restoring Pruned Large Language Models via Lost Component Compensation
- CPSVD: Enhancing Large Language Model Compression via Column-Preserving Singular Value Decomposition
- ARA: Adaptive Rank Allocation for Efficient Large Language Model SVD Compression
- From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
- Elastic ViTs from Pretrained Models without Retraining
- Mixed-Precision Quantization for Language Models: Techniques and Prospects
- Synera: Synergistic LLM Serving across Device and Cloud at Scale
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
- Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
- MosaicDiff: Training-free Structural Pruning for Diffusion Model Acceleration Reflecting Pretraining Dynamics
- ShishuLM: Lightweight Language Model with Hybrid Decoder-MLP Architecture and Paired Weight Sharing
- MC#: Mixture Compressor for Mixture-of-Experts Large Models
- Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization
- AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- Small is Sufficient: Reducing the World AI Energy Consumption Through Model Selection
- Accelerating Attention with Basis Decomposition
- AILoRA: Function-Aware Asymmetric Initialization for Low-Rank Adaptation of Large Language Models
- Fewer Weights, More Problems: A Practical Attack on LLM Pruning
- OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
- Vanishing Contributions: A Unified Framework for Smooth and Iterative Model Compression
- RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Models
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot
- lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
- Mixture of Neuron Experts
- Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
- ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization
- Expand Neurons, Not Parameters
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
- Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation
- CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
- Collaborative Compression for Large-Scale MoE Deployment on Edge
- Layer-wise dynamic rank for compressing large language models
- Effective Model Pruning
- UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
- A Second-Order Perspective on Pruning at Initialization and Knowledge Transfer
- Differentiable Sparsity via D-Gating: Simple and Versatile Structured Penalization
- DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
- PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
- Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
- Lightweight error mitigation strategies for post-training N:M activation sparsity in LLMs
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- StructPrune: Structured Global Pruning asymptotics with O(√(N)) GPU Memory
- RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
- Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs
- Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution
- WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning
- Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning
- GradMAP: Faster Layer Pruning with Gradient Metric and Projection Compensation
- Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
- Fair-GPTQ: Bias-Aware Quantization for Large Language Models
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
- Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
- Harnessing Optimization Dynamics for Curvature-Informed Model Merging
- GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
- Unified Start, Personalized End: Progressive Pruning for Efficient 3D Medical Image Segmentation
- DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
- COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
- Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
- Delta Activations: A Representation for Finetuned Large Language Models
- TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
- LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
- One VLM, Two Roles: Stage-Wise Routing and Specialty-Level Deployment for Clinical Workflows
- Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
- Towards On-Device Personalization: Cloud-device Collaborative Data Augmentation for Efficient On-device Language Model
- DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction
- Unraveling the cognitive patterns of Large Language Models through module communities
- GM-Skip: Metric-Guided Transformer Block Skipping for Efficient Vision-Language Models
- Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
- SparseMap: A Sparse Tensor Accelerator Framework Based on Evolution Strategy
- LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
- EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models
- READER: Retrieval-Assisted Drafter for Efficient LLM Inference
- P/D-Device: Disaggregated Large Language Model between Cloud and Devices
- LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
- Pushing the Envelope of LLM Inference on AI-PC
- Deep Language Geometry: Constructing a Metric Space from LLM Weights
- Pruning Large Language Models by Identifying and Preserving Functional Networks
- Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos
- S2M3: Split-and-Share Multi-Modal Models for Distributed Multi-Task Inference on the Edge
- CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
- Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
- LeanK: Learnable K Cache Channel Pruning for Efficient Decoding
- On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
- XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration via Multi-Head Speculative Decoding
- Unveiling Super Experts in Mixture-of-Experts Large Language Models
Discussions
Related