vix.ing · top · new · best · stats · spec

Gradient descent for deep equilibrium single-index models

2025/11/21 by Sanjit Dandapanthula, Aaditya Ramdas, Dandapanthula, Sanjit +1
Computer Science · Physics and Astronomy · #Advanced Graph Neural Networks #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Model Reduction and Neural Networks #Statistics Theory (math.ST) #Stochastic Gradient Optimization Techniques

paper · pdf · doi:10.48550/arxiv.2511.16976

openalex publication_date 2025/11/21 · openalex created_date 2025/11/25 · openalex updated_date 2026/07/30

Abstract

Deep equilibrium models (DEQs) have recently emerged as a powerful paradigm for training infinitely deep weight-tied neural networks that achieve state of the art performance across many modern machine learning tasks. Despite their practical success, theoretically understanding the gradient descent dynamics for training DEQs remains an area of active research. In this work, we rigorously study the gradient descent dynamics for DEQs in the simple setting of linear models and single-index models, filling several gaps in the literature. We prove a conservation law for linear DEQs which implies that the parameters remain trapped on spheres during training and use this property to show that gradient flow remains well-conditioned for all time. We then prove linear convergence of gradient descent to a global minimizer for linear DEQs and deep equilibrium single-index models under appropriate initialization and with a sufficiently small step size. Finally, we validate our theoretical findings through experiments.

Citations

Related