vix.ing · top · new · best · stats

Memorizing Gaussians with no over-parameterizaion via gradient decent on neural networks

2020/03/28 by Amit Daniely, Daniely, Amit · 7 citations
Computer Science · Mathematics · Psychology · #Artificial intelligence #Artificial neural network #Cognitive psychology #Computer science #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Memorization #Neural Networks and Applications #Psychology #Stochastic Gradient Optimization Techniques #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.2003.12895

published in arXiv (Cornell University) (Cornell University)

arxiv created 2020/03/28 · openalex publication_date 2020/03/28 · arxiv updated 2020/03/31 · openalex created_date 2020/04/03 · openalex updated_date 2026/07/28

Abstract

We prove that a single step of gradient decent over depth two network, with q hidden neurons, starting from orthogonal initialization, can memorize Ω((dq)/(log4(d))) independent and randomly labeled Gaussians in ℝd. The result is valid for a large class of activation functions, which includes the absolute value.

Citations

Related