2024/10/13 by Felix Benning, Leif Döring, Benning, Felix +1 · 1 citation
Computer Science · Engineering · Mathematics · #Algorithm #Combinatorics #Computer science #Dimension (graph theory) #Engineering #Fourth Dimension #Industrial Vision Systems and Defect Detection #Mathematics #Physics #Spacetime #Span (engineering) #Structural engineering #cs.LG #math.OC #math.PR #stat.ML
paper · pdf · doi:10.48550/arxiv.2410.09973
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/10/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We prove that all 'gradient span algorithms' have asymptotically deterministic behavior on scaled Gaussian random functions as the dimension tends to infinity. This is a functional generalization of similar results for random quadratic functions and spin glasses. They explain the counterintuitive phenomenon that different training runs of many large machine learning models result in approximately equal cost curves despite random initialization on a complicated non-convex landscape. This 'predictable progress' phenomenon is exploited by the AutoML community: Since the optimization progress of a single run is already representative, multiple retries with the same hyperparameters are not necessary.