2020/07/08 by Andrii Trelin, Trelin, Andrii, Aleš Procházka +1 · 2 citations
Computer Science · Engineering · Mathematics · #Arithmetic #Artificial intelligence #Binary number #Computer science #Data mining #Energy Load and Power Forecasting #FOS: Computer and information sciences #Feature (linguistics) #Feature selection #Linguistics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine learning #Mathematics #Neural Networks and Applications #Pattern recognition (psychology) #Philosophy #Selection (genetic algorithm) #Target Tracking and Data Fusion in Sensor Networks #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2007.03920
published in arXiv (Cornell University) (Cornell University)
arxiv created 2020/07/08 · openalex publication_date 2020/07/08 · arxiv updated 2020/07/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Feature selection is one of the most decisive tools in understanding data and machine learning models. Among other methods, sparsity induced by L1 penalty is one of the simplest and best studied approaches to this problem. Although such regularization is frequently used in neural networks to achieve sparsity of weights or unit activations, it is unclear how it can be employed in the feature selection problem. This work aims at extending the neural network with ability to automatically select features by rethinking how the sparsity regularization can be used, namely, by stochastically penalizing feature involvement instead of the layer weights. The proposed method has demonstrated superior efficiency when compared to a few classical methods, achieved with minimal or no computational overhead, and can be directly applied to any existing architecture. Furthermore, the method is easily generalizable for neuron pruning and selection of regions of importance for spectral data.