vix.ing · top · new · best · stats · spec

A Feature Selection Method that Controls the False Discovery Rate

2022/08/05 by Mehdi Rostami, Rostami, Mehdi, Olli Saarela +1
Computer Science · Mathematics · #FOS: Computer and information sciences #Imbalanced Data Classification Techniques #Machine Learning and Data Classification #Methodology (stat.ME) #Statistical Methods and Inference

paper · pdf · doi:10.48550/arxiv.2208.02948

openalex publication_date 2022/08/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The problem of selecting a handful of truly relevant variables in supervised machine learning algorithms is a challenging problem in terms of untestable assumptions that must hold and unavailability of theoretical assurances that selection errors are under control. We propose a distribution-free feature selection method, referred to as Data Splitting Selection (DSS) which controls False Discovery Rate (FDR) of feature selection while obtaining a high power. Another version of DSS is proposed with a higher power which "almost" controls FDR. No assumptions are made on the distribution of the response or on the joint distribution of the features. Extensive simulation is performed to compare the performance of the proposed methods with the existing ones.

Related