2022/09/02 by Jiahui Yu, Konstantinos Spiliopoulos, Yu, Jiahui +1 · 1 citation
Computer Science · Physics and Astronomy · #60F05 #60G99 #68T01 #Applications (stat.AP) #FOS: Computer and information sciences #FOS: Mathematics #Image and Signal Denoising Methods #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Model Reduction and Neural Networks #Neural Networks and Applications #Probability (math.PR)
paper · pdf · doi:10.48550/arxiv.2209.01018
openalex publication_date 2022/09/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We study the effect of normalization on the layers of deep neural networks of feed-forward type. A given layer i with Ni hidden units is allowed to be normalized by 1/Ni^γi with γi∈[1/2,1] and we study the effect of the choice of the γi on the statistical behavior of the neural network's output (such as variance) as well as on the test accuracy on the MNIST data set. We find that in terms of variance of the neural network's output and test accuracy the best choice is to choose the γi's to be equal to one, which is the mean-field scaling. We also find that this is particularly true for the outer layer, in that the neural network's behavior is more sensitive in the scaling of the outer layer as opposed to the scaling of the inner layers. The mechanism for the mathematical analysis is an asymptotic expansion for the neural network's output. An important practical consequence of the analysis is that it provides a systematic and mathematically informed way to choose the learning rate hyperparameters. Such a choice guarantees that the neural network behaves in a statistically robust way as the Ni grow to infinity.