Optimizing ML Training with Metagradient Descent
2025/03/17 by Logan Engstrom, Andrew Ilyas, Engstrom, Logan +9 · 12 voices · 11 citations
#stat.ML #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2503.13751
Abstract
A major challenge in training large-scale machine learning models is configuring the training process to maximize model performance, i.e., finding the best training setup from a vast design space. In this work, we unlock a gradient-based approach to this problem. We first introduce an algorithm for efficiently calculating metagradients -- gradients through model training -- at scale. We then introduce a "smooth model training" framework that enables effective optimization using metagradients. With metagradient descent (MGD), we greatly improve on existing dataset selection methods, outperform accuracy-degrading data poisoning attacks by an order of magnitude, and automatically find competitive learning rate schedules.
Cited by
Discussions
- Optimizing ML training with metagradient descent [hn, 83 points, 13 comments]
- arxiv.org/abs/2503.13751 [bsky, 1 points, 0 comments]
- https://bsky.app/profile/hackernews.com.web.brid.gy/post/3llal3iwsg4w2 [bsky, 0 points, 0 comments]
- Optimizing ML training with metagradient descent [bsky, 0 points, 0 comments]
- Optimizing ML training with metagradient descent #HackerNews https://arxiv.org/abs/2503.13751 [bsky, 0 points, 0 comments]
- Optimizing ML training with metagradient descent https://arxiv.org/abs/2503.13751 (https://news.ycombinator.com/item?id=43476134) [bsky, 0 points, 0 comments]
- Optimizing ML training with metagradient descent https://arxiv.org/abs/2503.13751 [comments] [71 points] [bsky, 0 points, 0 comments]
- Optimizing ML Training with Metagradient Descent https://arxiv.org/abs/2503.13751 [bsky, 0 points, 0 comments]
- ⚡ Hackernews Top story: Optimizing ML training with metagradient descent [bsky, 0 points, 0 comments]
- Optimizing ML training with metagradient descent (arxiv.org) Main Link | Discussion [bsky, 0 points, 0 comments]
- Optimizing ML Training with Metagradient Descent https://arxiv.org/abs/2503.13751 [bsky, 0 points, 0 comments]
- Optimizing ML training with metagradient descent https://arxiv.org/abs/2503.13751 (https://news.ycombinator.com/item?id=43476134) [bsky, 0 points, 0 comments]
Related