2019/02/14 by Anastasia Borovykh, Cornelis W. Oosterlee, Borovykh, Anastasia +3
Computer Science · #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural Networks and Applications #Time Series Analysis and Forecasting
paper · pdf · doi:10.48550/arxiv.1902.05312
openalex publication_date 2019/02/14 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
In this paper we study the generalization capabilities of fully-connected\nneural networks trained in the context of time series forecasting. Time series\ndo not satisfy the typical assumption in statistical learning theory of the\ndata being i.i.d. samples from some data-generating distribution. We use the\ninput and weight Hessians, that is the smoothness of the learned function with\nrespect to the input and the width of the minimum in weight space, to quantify\na network's ability to generalize to unseen data. While such generalization\nmetrics have been studied extensively in the i.i.d. setting of for example\nimage recognition, here we empirically validate their use in the task of time\nseries forecasting. Furthermore we discuss how one can control the\ngeneralization capability of the network by means of the training process using\nthe learning rate, batch size and the number of training iterations as\ncontrols. Using these hyperparameters one can efficiently control the\ncomplexity of the output function without imposing explicit constraints.\n