vix.ing · top · new · best · stats · spec

On Connecting Deep Trigonometric Networks with Deep Gaussian Processes:\n Covariance, Expressivity, and Neural Tangent Kernel

2022/03/14 by Chi-Ken Lu, Lu, Chi-Ken, Patrick Shafto +1 · 1 citation
Computer Science · Decision Sciences · #Gaussian Processes and Bayesian Inference #Forecasting Techniques and Applications

paper · pdf · doi:10.48550/arxiv.2203.07411

Abstract

Deep Gaussian Process (DGP) as a model prior in Bayesian learning intuitively\nexploits the expressive power in function composition. DGPs also offer diverse\nmodeling capabilities, but inference is challenging because marginalization in\nlatent function space is not tractable. With Bochner's theorem, DGP with\nsquared exponential kernel can be viewed as a deep trigonometric network\nconsisting of the random feature layers, sine and cosine activation units, and\nrandom weight layers. In the wide limit with a bottleneck, we show that the\nweight space view yields the same effective covariance functions which were\nobtained previously in function space. Also, varying the prior distributions\nover network parameters is equivalent to employing different kernels. As such,\nDGPs can be translated into the deep bottlenecked trig networks, with which the\nexact maximum a posteriori estimation can be obtained. Interestingly, the\nnetwork representation enables the study of DGP's neural tangent kernel, which\nmay also reveal the mean of the intractable predictive distribution.\nStatistically, unlike the shallow networks, deep networks of finite width have\ncovariance deviating from the limiting kernel, and the inner and outer widths\nmay play different roles in feature learning. Numerical simulations are present\nto support our findings.\n

Cited by

Related