2021/02/12 by Vincent Fortuin, Adrià Garriga-Alonso, Fortuin, Vincent +13 · 24 citations
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Artificial intelligence #Artificial neural network #Bayesian inference #Bayesian probability #Computer science #Convolutional neural network #FOS: Computer and information sciences #Gaussian #Gaussian Processes and Bayesian Inference #Inference #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification #Machine learning #Pattern recognition (psychology) #Prior probability #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2102.06571
published in arXiv (Cornell University) (Cornell University) · Accepted at ICLR 2022
openalex publication_date 2021/02/12 · arxiv created 2022/03/16 · arxiv updated 2022/03/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06
Isotropic Gaussian priors are the de facto standard for modern Bayesian neural network inference. However, it is unclear whether these priors accurately reflect our true beliefs about the weight distributions or give optimal performance. To find better priors, we study summary statistics of neural network weights in networks trained using stochastic gradient descent (SGD). We find that convolutional neural network (CNN) and ResNet weights display strong spatial correlations, while fully connected networks (FCNNs) display heavy-tailed weight distributions. We show that building these observations into priors can lead to improved performance on a variety of image classification datasets. Surprisingly, these priors mitigate the cold posterior effect in FCNNs, but slightly increase the cold posterior effect in ResNets.