vix.ing · top · new · best · stats · spec

The Disharmony between BN and ReLU Causes Gradient Explosion, but is Offset by the Correlation between Activations

2023/04/23 by Inyoung Paik, Jaesik Choi, Paik, Inyoung +1 · 1 citation
Computer Science · Neuroscience · #Brain Tumor Detection and Classification #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and ELM

paper · pdf · doi:10.48550/arxiv.2304.11692

openalex publication_date 2023/04/23 · openalex created_date 2023/04/27 · openalex updated_date 2026/07/28

Abstract

Deep neural networks, which employ batch normalization and ReLU-like activation functions, suffer from instability in the early stages of training due to the high gradient induced by temporal gradient explosion. In this study, we analyze the occurrence and mitigation of gradient explosion both theoretically and empirically, and discover that the correlation between activations plays a key role in preventing the gradient explosion from persisting throughout the training. Finally, based on our observations, we propose an improved adaptive learning rate algorithm to effectively control the training instability.

Cited by

Related