2020/10/05 by Aku Kammonen, Jonas Kiessling, Kammonen, Aku +9
Computer Science · Engineering · Mathematics · #65C05 #65D15 #65D40 #FOS: Mathematics #Geophysical Methods and Applications #Non-Destructive Testing Techniques #Numerical Analysis (math.NA) #Numerical methods in engineering #Optimization and Control (math.OC) #cs.NA #math.NA #math.OC #msc:65C05 #msc:65D15 #msc:65D40
paper · pdf · doi:10.48550/arxiv.2010.01887
openalex publication_date 2020/10/05 · arxiv created 2021/04/14 · arxiv updated 2021/04/15 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
Estimates of the generalization error are proved for a residual neural network with L random Fourier features layers zℓ+1= z_ℓ + Re∑k=1K bℓ ke^iωℓ k z_ℓ+ Re∑k=1K cℓ ke^iω'ℓ k⋅ x. An optimal distribution for the frequencies (ωℓ k,ω'ℓ k) of the random Fourier features e^iωℓ k z_ℓ and e^iω'ℓ k⋅ x is derived. This derivation is based on the corresponding generalization error for the approximation of the function values f(x). The generalization error turns out to be smaller than the estimate ‖ f‖2L1(ℝd)/(KL) of the generalization error for random Fourier features with one hidden layer and the same total number of nodes KL, in the case the L^∞-norm of f is much less than the L1-norm of its Fourier transform f. This understanding of an optimal distribution for random features is used to construct a new training method for a deep residual network. Promising performance of the proposed new algorithm is demonstrated in computational experiments.