vix.ing · top · new · best · stats

On the Neural Feature Ansatz for Deep Neural Networks

2025/10/17 by Edward Tansley, Tansley, Edward, Estelle Massart +3
Computer Science · Physics and Astronomy · #Ansatz #Artificial neural network #Counterexample #Deep learning #Exponent #FOS: Computer and information sciences #Feature (linguistics) #Generative Adversarial Networks and Image Synthesis #Initialization #Machine Learning (cs.LG) #Margin (machine learning) #Model Reduction and Neural Networks #Nonlinear system #Stochastic Gradient Optimization Techniques

paper · pdf · doi:10.48550/arxiv.2510.15563

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2025/10/17 · openalex created_date 2025/10/21 · openalex updated_date 2026/08/05

Abstract

Understanding feature learning is an important open question in establishing a mathematical foundation for deep neural networks. The Neural Feature Ansatz (NFA) states that after training, the Gram matrix of the first-layer weights of a deep neural network is proportional to some power α>0 of the average gradient outer product (AGOP) of this network with respect to its inputs. Assuming gradient flow dynamics with balanced weight initialization, the NFA was proven to hold throughout training for two-layer linear networks with exponent α= 1/2 (Radhakrishnan et al., 2024). We extend this result to networks with L ≥ 2 layers, showing that the NFA holds with exponent α= 1/L, thus demonstrating a depth dependency of the NFA. Furthermore, we prove that for unbalanced initialization, the NFA holds asymptotically through training if weight decay is applied. We also provide counterexamples showing that the NFA does not hold for some network architectures with nonlinear activations, even when these networks fit arbitrarily well the training data. We thoroughly validate our theoretical results through numerical experiments across a variety of optimization algorithms, weight decay rates and initialization schemes.

Citations

Related