vix.ing · top · new · best · stats · spec

Towards joint scaling laws with optimal batch size schedules

2026/07/30 by Jiaxiang Li, Zhiqi Bu, Shiyun Xu
Computer Science · Mathematics · #cs.LG #cs.AI #math.OC

paper · pdf

arxiv created 2026/07/30 · arxiv updated 2026/07/31

Abstract

Modern deep learning typically keeps the batch size static throughout training, thus overlooking the joint effect of learning rate and batch size on the training dynamics. In this paper, we study the deep learning dynamics through the lens of convex optimization and derive a joint characterization of loss in terms of both schedules, applicable to general optimizers and model architectures. This characterization yields a closed-form optimal batch size schedule for any prescribed learning rate schedule, and further leads to joint scaling laws that consistently outperform static batch size baselines, highlighting the significance of dynamic batch size schedule in large language model training.

Citations

Related