vix.ing · top · new · best · stats · spec

Statistically efficient thinning of a Markov chain sampler

2015/10/27 by Art B. Owen, Owen, Art B.
Engineering · #62M05 #65C40 #Computation (stat.CO) #FOS: Computer and information sciences #Industrial Vision Systems and Defect Detection #Machine Learning (cs.LG) #Machine Learning (stat.ML)

paper · pdf · doi:10.48550/arxiv.1510.07727

openalex publication_date 2015/10/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

It is common to subsample Markov chain output to reduce the storage burden. Geyer (1992) shows that discarding k-1 out of every k observations will not improve statistical efficiency, as quantified through variance in a given computational budget. That observation is often taken to mean that thinning MCMC output cannot improve statistical efficiency. Here we suppose that it costs one unit of time to advance a Markov chain and then θ>0 units of time to compute a sampled quantity of interest. For a thinned process, that cost θ is incurred less often, so it can be advanced through more stages. Here we provide examples to show that thinning will improve statistical efficiency if θ is large and the sample autocorrelations decay slowly enough. If the lag ℓ≥1 autocorrelations of a scalar measurement satisfy ρ_ℓ≥ρℓ+1≥0, then there is always a θ0 it is optimal if and only if θ≤ (1-ρ)2/(2ρ). This efficiency gain never exceeds 1+θ. This paper also gives efficiency bounds for autocorrelations bounded between those of two AR(1) processes.

Related