2025/05/21 by Sho Sonoda, Yuka Hashimoto, Sonoda, Sho +5
Economics, Econometrics and Finance · #Deep neural networks #Dynamical Systems (math.DS) #Entropy (arrow of time) #Exponential function #FOS: Computer and information sciences #FOS: Mathematics #Financial Markets and Investment Strategies #Generalization #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Metric (unit) #Semigroup #Upper and lower bounds #Variance (accounting)
paper · pdf · doi:10.48550/arxiv.2505.15064
published in ArXiv.org
openalex publication_date 2025/05/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Why and when does depth improve generalization? We study this question in an implementation-agnostic state-transition model, where a depth-k predictor is a readout class H composed with the word ball B(k,F) generated by hidden state transitions. Generalization bounds separate implementation error, approximation error, and statistical complexity, and upper bound the depth-dependent variance term by a Dudley entropy integral over B(k,F), with a conditional lower-bound diagnostic under readout separation. We identify geometric and semigroup mechanisms that keep this entropy contribution saturated or polynomial, and contrast them with separation mechanisms that recover the classical exponential-growth obstruction. Coupling these variance upper bounds with approximation rates gives typical depth trade-off patterns, clarifying that depth is statistically favorable when approximation improves rapidly while the transition semigroup remains geometrically tame.