vix.ing · top · new · best · stats · spec

Risk Bounds on MDL Estimators for Linear Regression Models with Application to Simple ReLU Neural Networks

2024/07/04 by Yoshinari Takeishi, Takeishi, Yoshinari, Jun’ichi Takeuchi +1 · 2 citations
Computer Science · #FOS: Computer and information sciences #Information Theory (cs.IT) #Neural Networks and Applications

paper · pdf · doi:10.48550/arxiv.2407.03854

openalex publication_date 2024/07/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

To investigate the theoretical foundations of deep learning from the viewpoint of the minimum description length (MDL) principle, we analyse risk bounds of MDL estimators based on two-stage codes for simple two-layers neural networks (NNs) with ReLU activation. For that purpose, we propose a method to design two-stage codes for linear regression models and establish an upper bound on the risk of the corresponding MDL estimators based on the theory of MDL estimators originated by Barron and Cover (1991). Then, we apply this result to the simple two-layers NNs with ReLU activation which consist of d nodes in the input layer, m nodes in the hidden layer and one output node. Since the object of estimation is only the m weights from the hidden layer to the output node in our setting, this is an example of linear regression models. As a result, we show that the redundancy of the obtained two-stage codes is small owing to the fact that the eigenvalue distribution of the Fisher information matrix of the NNs is strongly biased, which was shown by Takeishi et al. (2023) and has been refined in this paper. That is, we establish a tight upper bound on the risk of our MDL estimators. Note that our risk bound for the simple ReLU networks, of which the leading term is O(d2 log n /n), is independent of the number of parameters m.

Cited by

Related