vix.ing · top · new · best · stats · spec

Critical Percolation as a Framework to Analyze the Training of Deep Networks

2018/02/06 by Zohar Ringel, Ringel, Zohar, Rodrigo de +2
Computer Science · Mathematics · Physics and Astronomy · #Advanced Graph Neural Networks #Algorithm #Artificial intelligence #Combinatorics #Computer science #Deep learning #Disordered Systems and Neural Networks (cond-mat.dis-nn) #Euclidean geometry #Euclidean space #FOS: Computer and information sciences #FOS: Physical sciences #Focus (optics) #Function (biology) #Graph #Limit (mathematics) #Machine Learning (stat.ML) #Machine learning #Mathematics #Maxima and minima #Obstacle #Percolation (cognitive psychology) #Reachability #Statistical Mechanics (cond-mat.stat-mech) #Stochastic Gradient Optimization Techniques #Task (project management) #Theoretical computer science #Topological and Geometric Data Analysis #Topology (electrical circuits) #cond-mat.dis-nn #cond-mat.stat-mech #stat.ML

paper · pdf · doi:10.48550/arxiv.1802.02154

Accepted to ICLR 2018 as a conference paper

arxiv created 2018/02/06 · openalex publication_date 2018/02/06 · arxiv updated 2018/02/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In this paper we approach two relevant deep learning topics: i) tackling of graph structured input data and ii) a better understanding and analysis of deep networks and related learning algorithms. With this in mind we focus on the topological classification of reachability in a particular subset of planar graphs (Mazes). Doing so, we are able to model the topology of data while staying in Euclidean space, thus allowing its processing with standard CNN architectures. We suggest a suitable architecture for this problem and show that it can express a perfect solution to the classification task. The shape of the cost function around this solution is also derived and, remarkably, does not depend on the size of the maze in the large maze limit. Responsible for this behavior are rare events in the dataset which strongly regulate the shape of the cost function near this global minimum. We further identify an obstacle to learning in the form of poorly performing local minima in which the network chooses to ignore some of the inputs. We further support our claims with training experiments and numerical analysis of the cost function on networks with up to 128 layers.

Citations

Related