2021/09/01 by Ibrahim Alabdulmohsin, Hartmut Maennel, Alabdulmohsin, Ibrahim +3 · 4 citations
Computer Science · Mathematics · #68T07 #68T45 #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Algorithm #Artificial intelligence #Artificial neural network #Benchmark (surveying) #Computer science #Convolutional neural network #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generalization #Geography #Machine Learning (cs.LG) #Machine learning #Margin (machine learning) #Mathematics #Maxima and minima #Norm (philosophy) #Pattern recognition (psychology) #cs.LG #msc:68T07 #msc:68T45
paper · pdf · doi:10.48550/arxiv.2109.00267
published in arXiv (Cornell University) (Cornell University) · 12 figures, 7 tables
arxiv created 2021/09/01 · openalex publication_date 2021/09/01 · arxiv updated 2021/09/02 · openalex created_date 2021/09/13 · openalex updated_date 2026/07/28
Recent results suggest that reinitializing a subset of the parameters of a neural network during training can improve generalization, particularly for small training sets. We study the impact of different reinitialization methods in several convolutional architectures across 12 benchmark image classification datasets, analyzing their potential gains and highlighting limitations. We also introduce a new layerwise reinitialization algorithm that outperforms previous methods and suggest explanations of the observed improved generalization. First, we show that layerwise reinitialization increases the margin on the training examples without increasing the norm of the weights, hence leading to an improvement in margin-based generalization bounds for neural networks. Second, we demonstrate that it settles in flatter local minima of the loss surface. Third, it encourages learning general rules and discourages memorization by placing emphasis on the lower layers of the neural network. Our takeaway message is that the accuracy of convolutional neural networks can be improved for small datasets using bottom-up layerwise reinitialization, where the number of reinitialized layers may vary depending on the available compute budget.