vix.ing · top · new · best · stats

A variational approximate posterior for the deep Wishart process

2021/07/21 by Sebastian W. Ober, Laurence Aitchison, Ober, Sebastian W. +1
Computer Science · Mathematics · #Advanced Statistical Methods and Models #Algorithm #Applied mathematics #Artificial intelligence #Bayesian probability #Computer science #Determinantal point process #Discrete mathematics #Eigenvalues and eigenvectors #FOS: Computer and information sciences #Gaussian #Gaussian Processes and Bayesian Inference #Gaussian process #Inference #Kernel (algebra) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mathematical optimization #Mathematics #Positive-definite matrix #Prior probability #Random matrix #Statistical Methods and Bayesian Inference #Statistics #Wishart distribution #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.2107.10125

published in arXiv (Cornell University) 34 (Cornell University) · Accepted for publication at the 35th Conference on Neural Information Processing Systems (NeurIPS 2021). 23 pages

openalex publication_date 2021/07/21 · arxiv created 2021/12/03 · arxiv updated 2021/12/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/08

Abstract

Recent work introduced deep kernel processes as an entirely kernel-based alternative to NNs (Aitchison et al. 2020). Deep kernel processes flexibly learn good top-layer representations by alternately sampling the kernel from a distribution over positive semi-definite matrices and performing nonlinear transformations. A particular deep kernel process, the deep Wishart process (DWP), is of particular interest because its prior can be made equivalent to deep Gaussian process (DGP) priors for kernels that can be expressed entirely in terms of Gram matrices. However, inference in DWPs has not yet been possible due to the lack of sufficiently flexible distributions over positive semi-definite matrices. Here, we give a novel approach to obtaining flexible distributions over positive semi-definite matrices by generalising the Bartlett decomposition of the Wishart probability density. We use this new distribution to develop an approximate posterior for the DWP that includes dependency across layers. We develop a doubly-stochastic inducing-point inference scheme for the DWP and show experimentally that inference in the DWP can improve performance over doing inference in a DGP with the equivalent prior.

Citations

Related