2017/06/28 by Eugene Katsevich, Katsevich, Eugene, Chiara Sabatti +1 · 2 citations
Computer Science · #FOS: Computer and information sciences #Imbalanced Data Classification Techniques #Machine Learning and Data Classification #Methodology (stat.ME) #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1706.09375
openalex publication_date 2017/06/28 · openalex created_date 2022/10/02 · openalex updated_date 2026/07/28
We tackle the problem of selecting from among a large number of variables\nthose that are 'important' for an outcome. We consider situations where groups\nof variables are also of interest in their own right. For example, each\nvariable might be a genetic polymorphism and we might want to study how a trait\ndepends on variability in genes, segments of DNA that typically contain\nmultiple such polymorphisms. Or, variables might quantify various aspects of\nthe functioning of individual internet servers owned by a company, and we might\nbe interested in assessing the importance of each server as a whole on the\naverage download speed for the company's customers. In this context, to\ndiscover that a variable is relevant for the outcome implies discovering that\nthe larger entity it represents is also important. To guarantee meaningful and\nreproducible results, we suggest controlling the rate of false discoveries for\nfindings at the level of individual variables and at the level of groups.\nBuilding on the knockoff construction of Barber and Candes (2015) and the\nmultilayer testing framework of Barber and Ramdas (2016), we introduce the\nmultilayer knockoff filter (MKF). We prove that MKF simultaneously controls the\nFDR at each resolution and use simulations to show that it incurs little power\nloss compared to methods that provide guarantees only for the discoveries of\nindividual variables. We apply MKF to analyze a genetic dataset and find that\nit successfully reduces the number of false gene discoveries without a\nsignificant reduction in power.\n