vix.ing · top · new · best · stats · spec

FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization

2025/05/21 by Fangxin Liu, Liu, Fangxin, Zongwu Wang +11 · 2 citations
Computer Science · #Advanced Data Storage Technologies #Algorithms and Data Compression #FOS: Computer and information sciences #I.2.1 #I.2.7 #Machine Learning (cs.LG) #Neural Networks and Applications

paper · pdf · doi:10.48550/arxiv.2506.12024

openalex publication_date 2025/05/21 · openalex created_date 2025/10/13 · openalex updated_date 2026/07/28

Abstract

The rapid advancement of large language models (LLMs) has exacerbated the memory bottleneck due to the widening gap between model parameter scaling and hardware capabilities. While post-training quantization techniques effectively reduce memory overhead, existing methods predominantly rely on static quantization strategies, which struggle to adapt to dynamic workloads. To address this, we propose FlexQuant, a dynamic precision-switching framework that optimizes the trade-off between inference speed and accuracy. Leveraging model perplexity entropy and Kullback-Leibler divergence, FlexQuant enables fine-grained, layer-wise mixed-precision quantization and dynamically adjusts bit-widths during each token generation. FlexQuant provides a comprehensive analysis of quantization strategies, introduces a precision requirement model for optimal switching, and implements efficient fine-grained precision management. Evaluations demonstrate that FlexQuant achieves a 1.3x end-to-end speedup across diverse language tasks with negligible accuracy loss introduced. This framework offers a flexible and adaptive solution for efficient LLM deployment. Code is released at https://github.com/ZongwuWang/FlexQuant.git.

Cited by

Related