2020/08/13 by Guenther Walther, Walther, Guenther, Andrew B. Perry +2
Computer Science · Decision Sciences · Mathematics · Medicine · #Advanced Statistical Process Monitoring #Bayesian Methods and Mixture Models #Data-Driven Disease Surveillance #FOS: Mathematics #Statistical Methods and Inference #Statistics Theory (math.ST)
paper · pdf · doi:10.48550/arxiv.2008.06136
openalex publication_date 2020/08/13 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
We consider the problem of detecting an elevated mean on an interval with\nunknown location and length in the univariate Gaussian sequence model. Recent\nresults have shown that using scale-dependent critical values for the scan\nstatistic allows to attain asymptotically optimal detection simultaneously for\nall signal lengths, thereby improving on the traditional scan, but this\nprocedure has been criticized for losing too much power for short signals. We\nexplain this discrepancy by showing that these asymptotic optimality results\nwill necessarily be too imprecise to discern the performance of scan statistics\nin a practically relevant way, even in a large sample context. Instead, we\npropose to assess the performance with a new finite sample criterion. We then\npresent three calibrations for scan statistics that perform well across a range\nof relevant signal lengths: The first calibration uses a particular adjustment\nto the critical values and is therefore tailored to the Gaussian case. The\nsecond calibration uses a scale-dependent adjustment to the significance levels\nand is therefore applicable to arbitrary known null distributions. The third\ncalibration restricts the scan to a particular sparse subset of the scan\nwindows and then applies a weighted Bonferroni adjustment to the corresponding\ntest statistics. This sl Bonferroni scan is also applicable to arbitrary\nnull distributions and in addition is very simple to implement. We show how to\napply these calibrations for scanning in a number of distributional settings:\nfor normal observations with an unknown baseline and a known or unknown\nconstant variance,for observations from a natural exponential family, for\npotentially heteroscadastic observations from a symmetric density by employing\nself-normalization in a novel way, and for exchangeable observations using\ntests based on permutations, ranks or signs.\n