2018/01/08 by Daniel F. Schmidt, Enes Makalic, Schmidt, Daniel F. +2 · 2 citations
Computer Science · Mathematics · #Bayesian Methods and Mixture Models #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Statistical Methods and Bayesian Inference #Statistical Methods and Inference #Statistics Theory (math.ST)
paper · pdf · doi:10.48550/arxiv.1801.02321
openalex publication_date 2018/01/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Global-local shrinkage hierarchies are an important innovation in Bayesian\nestimation. We propose the use of log-scale distributions as a novel basis for\ngenerating familes of prior distributions for local shrinkage hyperparameters.\nBy varying the scale parameter one may vary the degree to which the prior\ndistribution promotes sparsity in the coefficient estimates. By examining the\nclass of distributions over the logarithm of the local shrinkage parameter that\nhave log-linear, or sub-log-linear tails, we show that many standard prior\ndistributions for local shrinkage parameters can be unified in terms of the\ntail behaviour and concentration properties of their corresponding marginal\ndistributions over the coefficients \βj. We derive upper bounds on the\nrate of concentration around |\βj|=0, and the tail decay as |\βj|\n\→ \∞, achievable by this wide class of prior distributions.\n We then propose a new type of ultra-heavy tailed prior, called the log-t\nprior with the property that, irrespective of the choice of associated scale\nparameter, the marginal distribution always diverges at \βj = 0, and\nalways possesses super-Cauchy tails. We develop results demonstrating when\nprior distributions with (sub)-log-linear tails attain Kullback--Leibler\nsuper-efficiency and prove that the log-t prior distribution is always\nsuper-efficient. We show that the log-t prior is less sensitive to\nmisspecification of the global shrinkage parameter than the horseshoe or lasso\npriors. By incorporating the scale parameter of the log-scale prior\ndistributions into the Bayesian hierarchy we derive novel adaptive shrinkage\nprocedures. Simulations show that the adaptive log-t procedure appears to\nalways perform well, irrespective of the level of sparsity or signal-to-noise\nratio of the underlying model.\n