vix.ing · top · new · best · stats · spec

The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent

2025/02/19 by Yatin Dandi, Luca Pesce, Dandi, Yatin +6 · 2 voices · 1 citation
Engineering · #3D Shape Modeling and Analysis #Industrial Vision Systems and Defect Detection

paper · pdf · doi:10.48550/arxiv.2502.13961

Abstract

Understanding the advantages of deep neural networks trained by gradient descent (GD) compared to shallow models remains an open theoretical challenge. In this paper, we introduce a class of target functions (single and multi-index Gaussian hierarchical targets) that incorporate a hierarchy of latent subspace dimensionalities. This framework enables us to analytically study the learning dynamics and generalization performance of deep networks compared to shallow ones in the high-dimensional limit. Specifically, our main theorem shows that feature learning with GD successively reduces the effective dimensionality, transforming a high-dimensional problem into a sequence of lower-dimensional ones. This enables learning the target function with drastically less samples than with shallow networks. While the results are proven in a controlled training setting, we also discuss more common training procedures and argue that they learn through the same mechanisms.

Cited by

Discussions

Related