vix.ing · top · new · best · stats

Tensor Programs IIb: Architectural Universality of Neural Tangent Kernel Training Dynamics

2021/05/08 by Greg Yang, Etai Littwin, Yang, Greg +1 · 8 citations
Computer Science · Mathematics · #Advanced Neural Network Applications #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #Probability (math.PR) #Stochastic Gradient Optimization Techniques #Tensor decomposition and applications #cs.LG #cs.NE #math.PR

paper · pdf · doi:10.48550/arxiv.2105.03703

ICML 2021

arxiv created 2021/05/08 · openalex publication_date 2021/05/08 · arxiv updated 2021/05/11 · openalex created_date 2021/05/24 · openalex updated_date 2026/07/28

Abstract

Yang (2020a) recently showed that the Neural Tangent Kernel (NTK) at initialization has an infinite-width limit for a large class of architectures including modern staples such as ResNet and Transformers. However, their analysis does not apply to training. Here, we show the same neural networks (in the so-called NTK parametrization) during training follow a kernel gradient descent dynamics in function space, where the kernel is the infinite-width NTK. This completes the proof of the *architectural universality* of NTK behavior. To achieve this result, we apply the Tensor Programs technique: Write the entire SGD dynamics inside a Tensor Program and analyze it via the Master Theorem. To facilitate this proof, we develop a graphical notation for Tensor Programs.

Citations

Cited by

Related