2006/02/01 by Ery Arias-Castro, David L. Donoho, Xiaoming Huo · 1 citation
Computer Science · Mathematics · #Geochemistry and Geologic Mapping #Medical Image Segmentation Techniques #Morphological variations and asymmetry #math.ST #msc:62G10 #msc:62G20 #msc:62M30 #stat.TH
paper · pdf · doi:10.1214/009053605000000787
published as Annals of Statistics 2006, Vol. 34, No. 1, 326-349 · Published at http://dx.doi.org/10.1214/009053605000000787 in the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)
openalex publication_date 2006/02/01 · arxiv created 2006/05/18 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We are given a set of n points that might be uniformly distributed in the unit square [0,1]2. We wish to test whether the set, although mostly consisting of uniformly scattered points, also contains a small fraction of points sampled from some (a priori unknown) curve with Cα-norm bounded by β. An asymptotic detection threshold exists in this problem; for a constant T−(α,β)>0, if the number of points sampled from the curve is smaller than T−(α,β)n1/(1+α), reliable detection is not possible for large n. We describe a multiscale significant-runs algorithm that can reliably detect concentration of data near a smooth curve, without knowing the smoothness information α or β in advance, provided that the number of points on the curve exceeds T*(α,β)n1/(1+α). This algorithm therefore has an optimal detection threshold, up to a factor T*/T−. At the heart of our approach is an analysis of the data by counting membership in multiscale multianisotropic strips. The strips will have area 2/n and exhibit a variety of lengths, orientations and anisotropies. The strips are partitioned into anisotropy classes; each class is organized as a directed graph whose vertices all are strips of the same anisotropy and whose edges link such strips to their “good continuations.” The point-cloud data are reduced to counts that measure membership in strips. Each anisotropy graph is reduced to a subgraph that consist of strips with significant counts. The algorithm rejects H0 whenever some such subgraph contains a path that connects many consecutive significant counts.