2023/06/29 by Subhadra Dasgupta, Holger Dette, Dasgupta, Subhadra +1
Biochemistry, Genetics and Molecular Biology · Mathematics · #FOS: Computer and information sciences #Gene expression and cancer classification #Methodology (stat.ME) #Statistical Methods and Inference
paper · pdf · doi:10.48550/arxiv.2306.16821
openalex publication_date 2023/06/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We propose a novel two-stage subsampling algorithm based on optimal design principles. In the first stage, we use a density-based clustering algorithm to identify an approximating design space for the predictors from an initial subsample. Next, we determine an optimal approximate design on this design space. Finally, we use matrix distances such as the Procrustes, Frobenius, and square-root distance to define the remaining subsample, such that its points are "closest" to the support points of the optimal design. Our approach reflects the specific nature of the information matrix as a weighted sum of non-negative definite Fisher information matrices evaluated at the design points and applies to a large class of regression models including models where the Fisher information is of rank larger than 1.