2019/06/18 by Valentin Thomas, Thomas, Valentin, Fabián Pedregosa +10 · 11 citations
Computer Science · Engineering · Mathematics · #Face and Expression Recognition #Sparse and Compressive Sensing Techniques #Stochastic Gradient Optimization Techniques #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1906.07774
Accepted to AISTATS 2020
arxiv created 2020/04/06 · arxiv updated 2020/04/08
The speed at which one can minimize an expected loss using stochastic methods depends on two properties: the curvature of the loss and the variance of the gradients. While most previous works focus on one or the other of these properties, we explore how their interaction affects optimization speed. Further, as the ultimate goal is good generalization performance, we clarify how both curvature and noise are relevant to properly estimate the generalization gap. Realizing that the limitations of some existing works stems from a confusion between these matrices, we also clarify the distinction between the Fisher matrix, the Hessian, and the covariance matrix of the gradients.