2015/05/03 by Rupesh K. Srivastava, Rupesh Kumar Srivastava, Klaus Greff +4 · 1 voice · 256 citations
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Business #Computer science #Stochastic Gradient Optimization Techniques #acm:68T01 #cs.LG #cs.NE #msc:68T01
paper · pdf · doi:10.48550/arxiv.1505.00387
published in arXiv (Cornell University) (Cornell University) · 6 pages, 2 figures. Presented at ICML 2015 Deep Learning workshop. Full paper is at arXiv:1507.06228
openalex publication_date 2015/05/03 · arxiv created 2015/11/03 · arxiv updated 2015/11/04 · openalex created_date 2024/04/10 · openalex updated_date 2026/08/05
There is plenty of theoretical and empirical evidence that depth of neural networks is a crucial ingredient for their success. However, network training becomes more difficult with increasing depth and training of very deep networks remains an open problem. In this extended abstract, we introduce a new architecture designed to ease gradient-based training of very deep networks. We refer to networks with this architecture as highway networks, since they allow unimpeded information flow across several layers on "information highways". The architecture is characterized by the use of gating units which learn to regulate the flow of information through a network. Highway networks with hundreds of layers can be trained directly using stochastic gradient descent and with a variety of activation functions, opening up the possibility of studying extremely deep and efficient architectures.