2025/05/13 by Jørgensen, Tollef Emil
#68T07 #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #I.2.0
paper · doi:10.48550/arxiv.2505.08620
Large language models have significantly advanced natural language processing, yet their heavy resource demands pose severe challenges regarding hardware accessibility and energy consumption. This paper presents a focused and high-level review of post-training quantization (PTQ) techniques designed to optimize the inference efficiency of LLMs by the end-user, including details on various quantization schemes, granularities, and trade-offs. The aim is to provide a balanced overview between the theory and applications of post-training quantization.