2019/01/13 by Zhixiang Xu, Xu, Zhixiang Eddie, Gao Huang +5 · 2 citations
Biochemistry, Genetics and Molecular Biology · #Cancer-related molecular mechanisms research #FOS: Computer and information sciences #Gene expression and cancer classification #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning in Bioinformatics
paper · pdf · doi:10.48550/arxiv.1901.04055
openalex publication_date 2019/01/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
A feature selection algorithm should ideally satisfy four conditions: reliably extract relevant features; be able to identify non-linear feature interactions; scale linearly with the number of features and dimensions; allow the incorporation of known sparsity structure. In this work we propose a novel feature selection algorithm, Gradient Boosted Feature Selection (GBFS), which satisfies all four of these requirements. The algorithm is flexible, scalable, and surprisingly straight-forward to implement as it is based on a modification of Gradient Boosted Trees. We evaluate GBFS on several real world data sets and show that it matches or out-performs other state of the art feature selection algorithms. Yet it scales to larger data set sizes and naturally allows for domain-specific side information.