SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
2023/01/02 by Elias Frantar, Dan Alistarh, Frantar, Elias +1 · 3 voices · 326 citations
Computer Science · Engineering · #Advanced Data Storage Technologies #Algorithm #Artificial intelligence #Code (set theory) #Computer science #Ferroelectric and Negative Capacitance Devices #Generative grammar #Generative model #Inference #Language model #Machine learning #One shot #Perplexity #Programming language #Pruning #Quantization (signal processing) #Source code #Topic Modeling #Transformer #cs.LG
paper · pdf · doi:10.48550/arxiv.2301.00774
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/01/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04
Abstract
We show for the first time that large-scale generative pretrained transformer (GPT) family models can be pruned to at least 50% sparsity in one-shot, without any retraining, at minimal loss of accuracy. This is achieved via a new pruning method called SparseGPT, specifically designed to work efficiently and accurately on massive GPT-family models. We can execute SparseGPT on the largest available open-source models, OPT-175B and BLOOM-176B, in under 4.5 hours, and can reach 60% unstructured sparsity with negligible increase in perplexity: remarkably, more than 100 billion weights from these models can be ignored at inference time. SparseGPT generalizes to semi-structured (2:4 and 4:8) patterns, and is compatible with weight quantization approaches. The code is available at: https://github.com/IST-DASLab/sparsegpt.
Cited by
- Unified Static-Dynamic Pruning for Efficient LLM Inference
- Importance-Aware OBS Pruning for Diffusion Models
- SHUFFLESPARSE: Learned Shuffles for Structured Sparse Networks
- GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding
- CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning
- Generalized Fisher-Weighted SVD: Scalable Kronecker-Factored Fisher Approximation for Compressing Large Language Models
- OrderMoE: An expert similarity driven distributed edge MoE inference
- It Takes a MAESTRO To Prune Bad Experts
- NIRVANA: Structured Pruning Reimagined for Large Language Model Compression
- Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models
- Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI
- Your Language Model Secretly Contains Personality Subnetworks
- Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression
- The 4/δ Bound: Designing Predictable LLM-Verifier Systems for Formal Method Guarantee
- 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
- Position: It's Time to Act on the Risk of Efficient Personalized Text Generation
- A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
- Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
- Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?
- TriSP: Tri-Signal Structured Pruning for Large Language Models
- Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices
- TimeBill: Time-Budgeted Inference for Large Language Models
- Pruning as a Game: Equilibrium-Driven Sparsification of Neural Networks
- Making Large Language Models Efficient Dense Retrievers
- Can abstract concepts from LLM improve SLM performance?
- MoE Pathfinder: Trajectory-driven Expert Pruning
- VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
- SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block Skipping
- Resting Neurons, Active Insights: Improving Input Sparsification for Large Language Models
- SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
- ZK-APEX: Zero-Knowledge Approximate Personalized Unlearning with Executable Proofs
- Leveraging KV Similarity for Online Structured Pruning in LLMs
- SQ-format: A Unified Sparse-Quantized Hardware-friendly Data Format for LLMs
- Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates
- Data-Free Pruning of Self-Attention Layers in LLMs
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
- Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
- HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
- SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot
- EfficientXpert: Efficient Domain Adaptation for Large Language Models via Propagation-Aware Pruning
- Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
- Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- INTERLACE: Interleaved Layer Pruning and Efficient Adaptation in Large Vision-Language Models
- ModHiFi: Identifying High Fidelity predictive components for Model Modification
- FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
- Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models
- Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
- Exploiting the Experts: Unauthorized Compression in MoE-LLMs
- E3-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
- Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers
- Automatic Pruning Discovery for Large Language Models
- PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
- TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZone
- Weight-sparse transformers have interpretable circuits
- MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
- Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
- KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
- A3: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures
- EcoSpa: Efficient Transformer Training with Coupled Sparsity
- APP: Accelerated Path Patching with Task-Specific Pruning
- TwIST: Rigging the Lottery in Transformers with Independent Subnetwork Training
- Optimal Singular Damage: Efficient LLM Inference in Low Storage Regimes
- IG-Pruning: Input-Guided Block Pruning for Large Language Models
- Continual Learning, Not Training: Online Adaptation For Agents
- AI Progress Should Be Measured by Capability-Per-Resource, Not Scale Alone: A Framework for Gradient-Guided Resource Allocation in LLMs
- 1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Models
- NeuronMLP: Efficient LLM Inference via Singular Value Decomposition Compression and Tiling on AWS Trainium
- FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference
- OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
- The Curious Case of In-Training Compression of State Space Models
- Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for On-Device LLMs
- PRO: Enabling Precise and Robust Text Watermark for Open-Source LLMs
- PAHQ: Accelerating Automated Circuit Discovery through Mixed-Precision Inference Optimization
- TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
- FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
- Frustratingly Easy Task-aware Pruning for Large Language Models
- The Structural Scalpel: Automated Contiguous Layer Pruning for Large Language Models
- Scaling Up Efficient Small Language Models Serving and Deployment for Semantic Job Search
- Restoring Pruned Large Language Models via Lost Component Compensation
- CPSVD: Enhancing Large Language Model Compression via Column-Preserving Singular Value Decomposition
- ARA: Adaptive Rank Allocation for Efficient Large Language Model SVD Compression
- From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
- Elastic ViTs from Pretrained Models without Retraining
- Mixed-Precision Quantization for Language Models: Techniques and Prospects
- Synera: Synergistic LLM Serving across Device and Cloud at Scale
- Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
- Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
- MosaicDiff: Training-free Structural Pruning for Diffusion Model Acceleration Reflecting Pretraining Dynamics
- ShishuLM: Lightweight Language Model with Hybrid Decoder-MLP Architecture and Paired Weight Sharing
- MC#: Mixture Compressor for Mixture-of-Experts Large Models
- Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization
- AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- Small is Sufficient: Reducing the World AI Energy Consumption Through Model Selection
- Accelerating Attention with Basis Decomposition
- AILoRA: Function-Aware Asymmetric Initialization for Low-Rank Adaptation of Large Language Models
- Fewer Weights, More Problems: A Practical Attack on LLM Pruning
- OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
- Vanishing Contributions: A Unified Framework for Smooth and Iterative Model Compression
- RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Models
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot
- lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
- Mixture of Neuron Experts
- Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
- ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization
- Expand Neurons, Not Parameters
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
- Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation
- CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
- Collaborative Compression for Large-Scale MoE Deployment on Edge
- Layer-wise dynamic rank for compressing large language models
- Effective Model Pruning
- UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
- A Second-Order Perspective on Pruning at Initialization and Knowledge Transfer
- Differentiable Sparsity via D-Gating: Simple and Versatile Structured Penalization
- DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
- PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
- Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
- Lightweight error mitigation strategies for post-training N:M activation sparsity in LLMs
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- StructPrune: Structured Global Pruning asymptotics with O(√(N)) GPU Memory
- RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
- Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs
- Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution
- WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning
- Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning
- GradMAP: Faster Layer Pruning with Gradient Metric and Projection Compensation
- Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
- Fair-GPTQ: Bias-Aware Quantization for Large Language Models
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
- Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
- Harnessing Optimization Dynamics for Curvature-Informed Model Merging
- GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
- Unified Start, Personalized End: Progressive Pruning for Efficient 3D Medical Image Segmentation
- DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
- COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
- Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
- Delta Activations: A Representation for Finetuned Large Language Models
- TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
- LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
- One VLM, Two Roles: Stage-Wise Routing and Specialty-Level Deployment for Clinical Workflows
- Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
- Towards On-Device Personalization: Cloud-device Collaborative Data Augmentation for Efficient On-device Language Model
- DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction
- Unraveling the cognitive patterns of Large Language Models through module communities
- GM-Skip: Metric-Guided Transformer Block Skipping for Efficient Vision-Language Models
- Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
- SparseMap: A Sparse Tensor Accelerator Framework Based on Evolution Strategy
- LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
- EGGS-PTP: An Expander-Graph Guided Structured Post-training Pruning Method for Large Language Models
- READER: Retrieval-Assisted Drafter for Efficient LLM Inference
- P/D-Device: Disaggregated Large Language Model between Cloud and Devices
- LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
- Pushing the Envelope of LLM Inference on AI-PC
- Deep Language Geometry: Constructing a Metric Space from LLM Weights
- Pruning Large Language Models by Identifying and Preserving Functional Networks
- Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos
- S2M3: Split-and-Share Multi-Modal Models for Distributed Multi-Task Inference on the Edge
- CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
- Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
- LeanK: Learnable K Cache Channel Pruning for Efficient Decoding
- On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
- XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration via Multi-Head Speculative Decoding
- Unveiling Super Experts in Mixture-of-Experts Large Language Models
- Detection Transformers Under the Knife: A Neuroscience-Inspired Approach to Ablations
- LoRA-PAR: A Flexible Dual-System LoRA Partitioning Approach to Efficient LLM Fine-Tuning
- Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
- Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method
- The Benchmark Illusion: Pruned LLMs Can Pass Multiple Choice but Fail to Answer
- DFQ-ViT: Data-Free Quantization for Vision Transformers without Fine-tuning
- On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention
- Toward Efficient SpMV in Sparse LLMs via Block Extraction and Compressed Storage
- FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich Training
- Beyond Bias Scores: Unmasking Vacuous Neutrality in Small Language Models
- Olica: Efficient Structured Pruning of Large Language Models without Retraining
- SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
- GeLaCo: An Evolutionary Approach to Layer Compression
- SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
- BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
- Compress Any Segment Anything Model (SAM)
- Boosting Parameter Efficiency in LLM-Based Recommendation through Sophisticated Pruning
- Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
- Unveiling Effective In-Context Configurations for Image Captioning: An External & Internal Analysis
- The Case for Instance-Optimized LLMs in OLAP Databases
- DOTResize: Reducing LLM Width via Discrete Optimal Transport-based Neuron Merging
- MGAA: Multi-Granular Adaptive Allocation fof Low-Rank Compression of LLMs
- BLaST: High Performance Inference and Pretraining using BLock Sparse Transformers
- From 2:4 to 8:16 sparsity patterns in LLMs for Outliers and Weights with Variance Correction
- High-Layer Attention Pruning with Rescaling
- Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization
- EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
- La RoSA: Enhancing LLM Efficiency via Layerwise Rotated Sparse Activation
- MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE
- Closed-Form Robustness Bounds for Second-Order Pruning of Neural Controller Policies
- Masked Gated Linear Unit
- Progtuning: Progressive Fine-tuning Framework for Transformer-based Language Models
- GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching
- DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
- Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models
- DipSVD: Dual-importance Protected SVD for Efficient LLM Compression
- Revisiting LoRA through the Lens of Parameter Redundancy: Spectral Encoding Helps
- SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity
- MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs
- LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling
- Training-free LLM Merging for Multi-task Learning
- Compression Aware Certified Training
- Learning Compact Vision Tokens for Efficient Large Multimodal Models
- SAFE: Finding Sparse and Flat Minima to Improve Pruning
- AWP: Activation-Aware Weight Pruning and Quantization with Projected Gradient Descent
- BAQ: Efficient Bit Allocation Quantization for Large Language Models
- Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias
- Kinetics: Rethinking Test-Time Scaling Laws
- Onboard Optimization and Learning: A Survey
- AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
- Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information
- SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling
- Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
- Faster MoE LLM Inference for Extremely Large Models
- CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
- Hopscotch: Discovering and Skipping Redundancies in Language Models
- FLEx: Personalized Federated Learning for Mixture-of-Experts LLMs via Expert Grafting
- FLoE: Fisher-Based Layer Selection for Efficient Sparse Adaptation of Low-Rank Experts
- SPAP: Structured Pruning via Alternating Optimization and Penalty Methods
- EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
- A Simple Linear Patch Revives Layer-Pruned Large Language Models
- Smooth Model Compression without Fine-Tuning
- TSENOR: Highly-Efficient Algorithm for Finding Transposable N:M Sparse Masks
- DenoiseRotator: Enhance Pruning Robustness for LLMs via Importance Concentration
- ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
- Leave it to the Specialist: Repair Sparse LLMs with Sparse Fine-Tuning via Sparsity Evolution
- LoKI: Low-damage Knowledge Implanting of Large Language Models
- ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
- SlimLLM: Accurate Structured Pruning for Large Language Models
- DLP: Dynamic Layerwise Pruning in Large Language Models
- TuneComp: Joint Fine-tuning and Compression for Large Foundation Models
- M-Wanda: Improving One-Shot Pruning for Multilingual LLMs
- Rethinking the Outlier Distribution in Large Language Models: An In-depth Study
- ALTER: All-in-One Layer Pruning and Temporal Expert Routing for Efficient Diffusion Generation
- LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions
- SHE-LoRA: Selective Homomorphic Encryption for Federated Tuning with Heterogeneous LoRA
- ResSVD: Residual Compensated SVD for Large Language Model Compression
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
- WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
- Pangu Light: Weight Re-Initialization for Pruning and Accelerating LLMs
- DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline Model
- Neural Parameter Search for Slimmer Fine-Tuned Models and Better Transfer
- μ-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
- Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding
- LatentLLM: Attention-Aware Joint Tensor Compression
- How Many Parameters Does Your Task Really Need? Task Specific Pruning with LLM-Sieve
- Two-Stage Regularization-Based Structured Pruning for LLMs
- TRIM: Achieving Extreme Sparsity with Targeted Row-wise Iterative Metric-driven Pruning
- KNN-SSD: Enabling Dynamic Self-Speculative Decoding via Nearest Neighbor Layer Set Optimization
- LLM Fingerprinting via Semantically Conditioned Watermarks
- LLM-Powered AI Agent Systems and Their Applications in Industry
- Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse
- Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models
- RAP: Runtime Adaptive Pruning for LLM Inference
- Boost Post-Training Quantization via Null Space Optimization for Large Language Models
- Improved Methods for Model Pruning and Knowledge Distillation
- Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization
- CE-LSLM: Efficient Large-Small Language Model Inference and Communication via Cloud-Edge Collaboration
- Exploring Federated Pruning for Large Language Models
- Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
- One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models
- Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform
- Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
- AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction
- F-WANDA: Fisher-Reweighted Post-Training Pruning for Sustainable Deployment of Large Language Models
- Accurate KV Cache Quantization with Outlier Tokens Tracing
- Addition is almost all you need: Compressing large language models with double binary factorization
- A Survey on Deep Neural Network Pruning: Taxonomy, Comparison, Analysis, and Recommendations
- Constrained Edge AI Deployment: Fine-Tuning vs Distillation for LLM Compression
- Semantic Retention and Extreme Compression in LLMs: Can We Have Both?
- GPrune-LLM: Generalization-Aware Structured Pruning for Large Language Models
- A Hamiltonian-Inspired Local-Operator Ansatz for Slimming Large Language Models
- FloE: On-the-Fly MoE Inference on Memory-constrained GPU
- Scalable LLM Math Reasoning Acceleration with Low-rank Distillation
- RAP: KV-Cache Compression via RoPE-Aligned Pruning
- Optimization over Trained (and Sparse) Neural Networks: A Surrogate within a Surrogate
- Efficient Shapley Value-based Non-Uniform Pruning of Large Language Models
- Efficient Fine-Tuning of Quantized Models via Adaptive Rank and Bitwidth
- Position: Enough of Scaling LLMs! Lets Focus on Downscaling
- On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
- TENP: Trapezoidal Expert Neuron Pruning For Mixture-of-Experts
- Efficient LLMs with AMP: Attention Heads and MLP Pruning
- Legilimens: Performant Video Analytics on the System-on-Chip Edge
- The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning
- Fast Entropy Decoding for Sparse MVM on GPUs
- Routing the Lottery: Adaptive Subnetworks for Heterogeneous Data
- R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
- Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory
- M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
- Beyond Peak TOPS/W: A System-Level Perspective on Hybrid Digital, Analogue and Neuromorphic Computing
- Omega-S: A Functional Resilience Index for LLM Fine-Tuning
- BrAIcht, a theatrical agent that speaks like Bertolt Brecht's characters
- When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning
- L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference
- The Rise of Small Language Models in Healthcare: A Comprehensive Survey
- ConTextual: Improving Clinical Text Summarization in LLMs with Context-preserving Token Filtering and Knowledge Graphs
- Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
- Compression Laws for Large Language Models
- Saliency-driven Dynamic Token Pruning for Large Language Models
- NoWag: A Unified Framework for Shape Preserving Compression of Large Language Models
- Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
- From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs
- Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
- Collaborative Learning of On-Device Small Model and Cloud-Based Large Model: Advances and Future Directions
- APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning
- Effective pruning of task-trained recurrent neural networks using noisy fluctuations and connection rescaling
- Persona-Pruner: Sculpting Lightweight Models for Role-Playing
- HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving
- TAMP: Token-Adaptive Layerwise Pruning in Multimodal Large Language Models
- SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
- Scaling Up On-Device LLMs via Active-Weight Swapping Between DRAM and Flash
- SD2: Self-Distilled Sparse Drafters
- Mosaic: Composite Projection Pruning for Resource-efficient LLMs
- Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
Discussions
Related