vix.ing · top · new · best · stats · spec

On the interplay of network structure and gradient convergence in deep\n learning

2015/11/17 by Ithapu, Vamsi K, Ravi, Sathya N, Singh Vikas +1
Engineering · Computer Science · #Sparse and Compressive Sensing Techniques #Stochastic Gradient Optimization Techniques #Domain Adaptation and Few-Shot Learning

paper · pdf · doi:10.48550/arxiv.1511.05297

Abstract

The regularization and output consistency behavior of dropout and layer-wise\npretraining for learning deep networks have been fairly well studied. However,\nour understanding of how the asymptotic convergence of backpropagation in deep\narchitectures is related to the structural properties of the network and other\ndesign choices (like denoising and dropout rate) is less clear at this time. An\ninteresting question one may ask is whether the network architecture and input\ndata statistics may guide the choices of learning parameters and vice versa. In\nthis work, we explore the association between such structural, distributional\nand learnability aspects vis- `a-vis their interaction with parameter\nconvergence rates. We present a framework to address these questions based on\nconvergence of backpropagation for general nonconvex objectives using\nfirst-order information. This analysis suggests an interesting relationship\nbetween feature denoising and dropout. Building upon these results, we obtain a\nsetup that provides systematic guidance regarding the choice of learning\nparameters and network sizes that achieve a certain level of convergence (in\nthe optimization sense) often mediated by statistical attributes of the inputs.\nOur results are supported by a set of experimental evaluations as well as\nindependent empirical observations reported by other groups.\n

Related