2021/12/31 by Inbar Seroussi, Seroussi, Inbar, Gadi Naveh +3 · 5 citations
Computer Science · Materials Science · #Data Analysis #FOS: Computer and information sciences #FOS: Physical sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning in Materials Science #Neural Networks and Applications #Statistics and Probability (physics.data-an) #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.2112.15383
openalex publication_date 2021/12/31 · openalex created_date 2022/05/05 · openalex updated_date 2026/07/28
Deep neural networks (DNNs) are powerful tools for compressing and distilling information. Their scale and complexity, often involving billions of inter-dependent parameters, render direct microscopic analysis difficult. Under such circumstances, a common strategy is to identify slow variables that average the erratic behavior of the fast microscopic variables. Here, we identify a similar separation of scales occurring in fully trained finitely over-parameterized deep convolutional neural networks (CNNs) and fully connected networks (FCNs). Specifically, we show that DNN layers couple only through the second moment (kernels) of their activations and pre-activations. Moreover, the latter fluctuates in a nearly Gaussian manner. For infinite width DNNs, these kernels are inert, while for finite ones they adapt to the data and yield a tractable data-aware Gaussian Process. The resulting thermodynamic theory of deep learning yields accurate predictions in various settings. In addition, it provides new ways of analyzing and understanding DNNs in general.