2025/03/19 by Jin, Lisa, Jianhao Ma, Zechun Liu +8 · 4 citations
Medicine · #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Medical Imaging Techniques and Applications #Optimization and Control (math.OC)
paper · pdf · doi:10.48550/arxiv.2503.15748
openalex publication_date 2025/03/19 · openalex created_date 2025/10/17 · openalex updated_date 2026/07/28
We develop a principled method for quantization-aware training (QAT) of large-scale machine learning models. Specifically, we show that convex, piecewise-affine regularization (PAR) can effectively induce the model parameters to cluster towards discrete values. We minimize PAR-regularized loss functions using an aggregate proximal stochastic gradient method (AProx) and prove that it has last-iterate convergence. Our approach provides an interpretation of the straight-through estimator (STE), a widely used heuristic for QAT, as the asymptotic form of PARQ. We conduct experiments to demonstrate that PARQ obtains competitive performance on convolution- and transformer-based vision tasks.