Mixed-Precision Quantization for Language Models: Techniques and Prospects
2025/10/19 by Mariam Rakka, Marios Fournarakis, Rakka, Mariam +13
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.2510.16805
Abstract
The rapid scaling of language models (LMs) has resulted in unprecedented computational, memory, and energy requirements, making their training and deployment increasingly unsustainable. Quantization has emerged as an essential compression technique to reduce model size, alleviate memory bottlenecks, and accelerate inference. However, while uniform low-bit quantization (e.g., INT8, INT4) provides significant efficiency gains, it can degrade accuracy in sensitive components of transformer-based LMs. Mixed-precision quantization offers a promising alternative by selectively allocating precision across layers or within tensors to balance efficiency and accuracy. This survey provides a comprehensive overview of Mixed-Precision quantization frameworks for LMs (MXPLMs). We first review quantization fundamentals, including uniform and non-uniform quantizers, quantization granularity, and methods widely used in post-training quantization. We then categorize and compare recent MXPLM frameworks according to their bit allocation strategies and precision configurations across weights, activations, and key-value caches. A comparative analysis highlights differences in perplexity, zero-shot task performance, and deployment trade-offs. Furthermore, we contrast MXPLMs with earlier mixed-precision quantization methods for deep neural networks, identifying strategies that transfer and those that face challenges in the LM setting. Finally, we summarize open issues and future directions, including hardware-aware design, activation quantization, and scalable optimization methods for billion-parameter models. By consolidating recent advances, this work serves as a reference for understanding the current landscape and research prospects of mixed-precision quantization for large-scale language models.
Citations
- Pretraining Large Language Models with NVFP4
- AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
- Sensitivity-Aware Post-Training Quantization for Deep Neural Networks
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models
- PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
- Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
- Gemma 3 Technical Report
- Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
- MergeQuant: Accurate 4-bit Static Quantization of Large Language Models by Channel-wise Calibration
- Towards Reasoning Ability of Small Language Models
- Optimizing Large Language Model Training Using FP4 Quantization
- BeST -- A Novel Source Selection Metric for Transfer Learning
- Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
- MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design
- ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
- FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
- BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
- A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness
- HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
- BF-IMNA: A Bit Fluid In-Memory Neural Architecture for Neural Network Acceleration
- A Comprehensive Study on Quantization Techniques for Large Language Models
- Progressive Mixed-Precision Decoding for Efficient LLM Inference
- Harnessing Your DRAM and SSD for Sustainable and Accessible LLM Inference with Mixed-Precision and Multi-level Caching
- Channel-Wise Mixed-Precision Quantization for Large Language Models
- A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms
- ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models
- LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
- EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
- Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models
- Low-Rank Quantization-Aware Training for LLMs
- DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
- To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
- SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
- Layer-Condensed KV Cache for Efficient Inference of Large Language Models
- QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
- Collage: Light-Weight Low-Precision Strategy for LLM Training
- Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
- A Survey on Efficient Inference for Large Language Models
- Cherry on Top: Parameter Heterogeneity and Quantization in Large Language Models
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
- LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
- Evaluating Quantized Large Language Models
- No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
- FlattenQuant: Breaking Through the Inference Compute-bound for Large Language Models with Per-tensor Quantization
- Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
- A Survey on Transformer Compression
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
- Efficient Large Language Models: A Survey
- Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
- Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
- Efficient LLM Inference on CPUs
- Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
- QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models
- PB-LLM: Partially Binarized Large Language Models
- Training and inference of large language models using 8-bit floating point
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
- A Survey on Model Compression for Large Language Models
- FLIQS: One-Shot Mixed-Precision Floating-Point and Integer Quantization Search
- A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
- Zero-Shot Neural Architecture Search: Challenges, Solutions, and Opportunities
- SqueezeLLM: Dense-and-Sparse Quantization
- SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
- OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
- FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
- LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
- QLoRA: Efficient Finetuning of Quantized LLMs
- Memory-Efficient Fine-Tuning of Compressed Large Language Models via sub-4-bit Integer Quantization
- Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond
- Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling
- TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings
- RPTQ: Reorder-based Post-training Quantization for Large Language Models
- GPT-4 Technical Report
- SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- FP8 Formats for Deep Learning
- Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and Pruning
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers
- PaLM: Scaling Language Modeling with Pathways
- Training Compute-Optimal Large Language Models
- Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
- Efficient Softmax Approximation for Deep Neural Networks with Attention Mechanism
- A White Paper on Neural Network Quantization
- VS-Quant: Per-vector Scaled Quantization for Accurate Low-Precision Neural Network Inference
- Sparsity in Deep Learning: Pruning and growth for efficient inference\n and training in neural networks
- Language Models are Few-Shot Learners
- Rethinking Differentiable Search for Mixed-Precision Neural Networks
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Bit Fusion: Bit-Level Dynamically Composable Architecture for\n Accelerating Deep Neural Networks
- Attention Is All You Need
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Cited by
Related