2020/08/08 by Ilona Kulikovskikh, Kulikovskikh, Ilona, Tarzan Legović +1
Computer Science · Mathematics · #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Markov Chains and Monte Carlo Methods #Neural and Evolutionary Computing (cs.NE) #Stochastic Gradient Optimization Techniques #cs.LG #cs.NE #stat.ML
paper · pdf · doi:10.48550/arxiv.2008.03501
arxiv created 2020/08/08 · openalex publication_date 2020/08/08 · arxiv updated 2020/08/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Current expectations from training deep learning models with gradient-based methods include: 1) transparency; 2) high convergence rates; 3) high inductive biases. While the state-of-art methods with adaptive learning rate schedules are fast, they still fail to meet the other two requirements. We suggest reconsidering neural network models in terms of single-species population dynamics where adaptation comes naturally from open-ended processes of "growth" and "harvesting". We show that the stochastic gradient descent (SGD) with two balanced pre-defined values of per capita growth and harvesting rates outperform the most common adaptive gradient methods in all of the three requirements.