2024/11/06 by K. Akiyama, Akiyama, Keito
Computer Science · Physics and Astronomy · #35Q49 #49J20 #82C32 #Analysis of PDEs (math.AP) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Model Reduction and Neural Networks #Neural Networks and Applications #Statistical Mechanics and Entropy
paper · pdf · doi:10.48550/arxiv.2411.03611
openalex publication_date 2024/11/06 · openalex created_date 2024/11/15 · openalex updated_date 2026/07/28
In recent years, learning for neural networks can be viewed as optimization in the space of probability measures. To obtain the exponential convergence to the optimizer, the regularizing term based on Shannon entropy plays an important role. Even though an entropy function heavily affects convergence results, there is almost no result on its generalization, because of the following two technical difficulties: one is the lack of sufficient condition for generalized logarithmic Sobolev inequality, and the other is the distributional dependence of the potential function within the gradient flow equation. In this paper, we establish a framework that utilizes a linearized potential function via Csiszár type of Tsallis entropy, which is one of the generalized entropies. We also show that our new framework enable us to derive an exponential convergence result.