vix.ing · top · new · best · stats · spec

Needles and straw in haystacks: Empirical Bayes estimates of possibly sparse sequences

2004/08/01 by Iain M. Johnstone, Bernard W. Silverman · 13 citations
Computer Science · Mathematics · #Bayesian Methods and Mixture Models #Blind Source Separation Techniques #Image and Signal Denoising Methods #math.ST #msc:62C12 #msc:62G05 #stat.TH

paper · pdf · doi:10.1214/009053604000000030

published as Annals of Statistics 2004, Vol. 32, No. 4, 1594-1649 · Published by the Institute of Mathematical Statistics (http://www.imstat.org) in the Annals of Statistics (http://www.imstat.org/aos/) at http://dx.doi.org/10.1214/009053604000000030

openalex publication_date 2004/08/01 · arxiv created 2004/10/05 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

An empirical Bayes approach to the estimation of possibly sparse sequences observed in Gaussian white noise is set out and investigated. The prior considered is a mixture of an atom of probability at zero and a heavy-tailed density γ, with the mixing weight chosen by marginal maximum likelihood, in the hope of adapting between sparse and dense sequences. If estimation is then carried out using the posterior median, this is a random thresholding procedure. Other thresholding rules employing the same threshold can also be used. Probability bounds on the threshold chosen by the marginal maximum likelihood approach lead to overall risk bounds over classes of signal sequences of length n, allowing for sparsity of various kinds and degrees. The signal classes considered are “nearly black” sequences where only a proportion η is allowed to be nonzero, and sequences with normalized ℓp norm bounded by η, for η>0 and 0<p≤2. Estimation error is measured by mean qth power loss, for 0<q≤2. For all the classes considered, and for all q in (0,2], the method achieves the optimal estimation rate as n→∞ and η→0 at various rates, and in this sense adapts automatically to the sparseness or otherwise of the underlying signal. In addition the risk is uniformly bounded over all signals. If the posterior mean is used as the estimator, the results still hold for q>1. Simulations show excellent performance. For appropriately chosen functions γ, the method is computationally tractable and software is available. The extension to a modified thresholding method relevant to the estimation of very sparse sequences is also considered.

Citations

Cited by