2019/02/02 by Sho Sonoda, Sonoda, Sho
Computer Science · Engineering · #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Sparse and Compressive Sensing Techniques #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.1902.00648
openalex publication_date 2019/02/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
An infinitely wide model is a weighted integration \∫ \φ(x,v) d\n\μ(v) of feature maps. This model excels at handling an infinite number of\nfeatures, and thus it has been adopted to the theoretical study of deep\nlearning. Kernel quadrature is a kernel-based numerical integration scheme\ndeveloped for fast approximation of expectations \∫ f(x) d p(x). In this\nstudy, regarding the weight \μ as a signed (or complex/vector-valued)\ndistribution of parameters, we develop the general kernel quadrature (GKQ) for\nparameter distributions. The proposed method can achieve a fast approximation\nrate O(e-p) with parameter number p, which is faster than the\ntraditional Barron's rate, and a fast estimation rate widetildeO(1/n) with\nsample size n. As a result, we have obtained a new norm-based complexity\nmeasure for infinitely wide models. Since the GKQ implicitly conducts the\nempirical risk minimization, we can understand that the complexity measure also\nreflects the generalization performance in the gradient learning setup.\n