2023/07/24 by Anav Sood, Trevor Hastie, Sood, Anav +1 · 2 citations
Computer Science · Mathematics · #Artificial intelligence #Bayesian Modeling and Causal Inference #Column (typography) #Computer science #Curse of dimensionality #Data Mining Algorithms and Applications #Data Structures and Algorithms (cs.DS) #Data mining #Data set #Dimensionality reduction #FOS: Computer and information sciences #Feature selection #Machine Learning (cs.LG) #Machine Learning and Data Classification #Mathematics #Methodology (stat.ME) #Parametric statistics #Selection (genetic algorithm) #Set (abstract data type) #Statistical hypothesis testing #Statistics
paper · pdf · doi:10.48550/arxiv.2307.12892
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/07/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We consider the problem of selecting a small subset of representative variables from a large dataset. In the computer science literature, this dimensionality reduction problem is typically formalized as Column Subset Selection (CSS). Meanwhile, the typical statistical formalization is to find an information-maximizing set of Principal Variables. This paper shows that these two approaches are equivalent, and moreover, both can be viewed as maximum likelihood estimation within a certain semi-parametric model. Within this model, we establish suitable conditions under which the CSS estimate is consistent in high dimensions, specifically in the proportional asymptotic regime where the number of variables over the sample size converges to a constant. Using these connections, we show how to efficiently (1) perform CSS using only summary statistics from the original dataset; (2) perform CSS in the presence of missing and/or censored data; and (3) select the subset size for CSS in a hypothesis testing framework.