2022/02/01 by Kris Sankaran, Sankaran, Kris
Biochemistry, Genetics and Molecular Biology · Computer Science · #Computation (stat.CO) #Data Analysis with R #FOS: Computer and information sciences #Gene expression and cancer classification
paper · pdf · doi:10.48550/arxiv.2202.00180
openalex publication_date 2022/02/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Algorithmic feature learners provide high-dimensional vector representations for non-matrix structured data, like image or text collections. Low-dimensional projections derived from these representations, called embeddings, are often used to explore variation in these data. However, it is not clear how to assess the embedding uncertainty. We adapt methods developed for bootstrapping principal components analysis to the setting where features are algorithmically derived from nonmatrix data. We empirically compare the derived confidence areas in simulations, varying factors influencing feature learning and the bootstrap, like feature learning algorithm complexity and bootstrap sample size. We illustrate the proposed approaches on a spatial proteomics dataset, where we observe that embedding precision is not uniform across all tissue types. Code, data, and pretrained models are available as an R compendium in the supplementary materials. Supplementary files for this article are available online.