vix.ing · top · new · best · stats · spec

Abide by the Law and Follow the Flow: Conservation Laws for Gradient\n Flows

2023/06/30 by Sibylle Marcotte, Rémi Gribonval, Marcotte, Sibylle +3 · 1 voice · 8 citations
Computer Science · Engineering · Mathematics · Physics and Astronomy · #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Mathematics #Fluid Dynamics and Turbulent Flows #Machine Learning (cs.LG) #Model Reduction and Neural Networks #Optimization and Control (math.OC) #Stochastic Gradient Optimization Techniques #cs.LG #math.OC

paper · pdf · doi:10.48550/arxiv.2307.00144

openalex publication_date 2023/06/30 · arxiv published 2023/06/30 · arxiv updated 2024/07/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Understanding the geometric properties of gradient descent dynamics is a key\ningredient in deciphering the recent success of very large machine learning\nmodels. A striking observation is that trained over-parameterized models retain\nsome properties of the optimization initialization. This "implicit bias" is\nbelieved to be responsible for some favorable properties of the trained models\nand could explain their good generalization properties. The purpose of this\narticle is threefold. First, we rigorously expose the definition and basic\nproperties of "conservation laws", that define quantities conserved during\ngradient flows of a given model (e.g. of a ReLU network with a given\narchitecture) with any training data and any loss. Then we explain how to find\nthe maximal number of independent conservation laws by performing\nfinite-dimensional algebraic manipulations on the Lie algebra generated by the\nJacobian of the model. Finally, we provide algorithms to: a) compute a family\nof polynomial laws; b) compute the maximal number of (not necessarily\npolynomial) independent conservation laws. We provide showcase examples that we\nfully work out theoretically. Besides, applying the two algorithms confirms for\na number of ReLU network architectures that all known laws are recovered by the\nalgorithm, and that there are no other independent laws. Such computational\ntools pave the way to understanding desirable properties of optimization\ninitialization in large machine learning models.\n

Citations

Cited by

Discussions

Related