2021/12/13 by Dimitris N. Politis, Politis, Dimitris N. · 2 citations
Computer Science · Decision Sciences · Mathematics · #Advanced Statistical Process Monitoring #FOS: Mathematics #Machine Learning and Algorithms #Statistical Methods and Inference #Statistics Theory (math.ST)
paper · pdf · doi:10.48550/arxiv.2112.06434
openalex publication_date 2021/12/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Subsampling is a general statistical method developed in the 1990s aimed at estimating the sampling distribution of a statistic θn in order to conduct nonparametric inference such as the construction of confidence intervals and hypothesis tests. Subsampling has seen a resurgence in the Big Data era where the standard, full-resample size bootstrap can be infeasible to compute. Nevertheless, even choosing a single random subsample of size b can be computationally challenging with both b and the sample size n being very large. In the paper at hand, we show how a set of appropriately chosen, non-random subsamples can be used to conduct effective -- and computationally feasible -- distribution estimation via subsampling. Further, we show how the same set of subsamples can be used to yield a procedure for subsampling aggregation -- also known as subagging -- that is scalable with big data. Interestingly, the scalable subagging estimator can be tuned to have the same (or better) rate of convergence as compared to θn. The paper is concluded by showing how to conduct inference, e.g., confidence intervals, based on the scalable subagging estimator instead of the original θn.