Beomseok Kwon
- LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
2022/06/20 by Gunho Park, Park, Gunho, Baeseong Park +18 · 1 voice · 26 citations
Engineering · Computer Science · #cs.DC #cs.CL
- No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
2024/02/28 by June Yong Yang, Yang, June Yong, Byeongwook Kim +13 · 20 citations
Computer Science · #Advanced Data Compression Techniques #Algorithms and Data Compression #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Medical Image Segmentation Techniques