vix.ing · top · new · best · stats

Stochastic mutual information gradient estimation for dimensionality reduction networks

2021/04/20 by Ozan Özdenizci, Ozan Ozdenizci, Deniz Erdogmus +1 · 20 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Artificial intelligence #Artificial neural network #Class (philosophy) #Computer science #Curse of dimensionality #Data mining #Dimensionality reduction #Discriminative model #Face and Expression Recognition #Feature (linguistics) #Feature selection #Feature vector #Gene expression and cancer classification #Machine Learning and Data Classification #Machine learning #Mathematics #Mutual information #Pattern recognition (psychology) #Ranking (information retrieval) #Reduction (mathematics) #cs.IT #cs.LG #math.IT #stat.ML

paper · pdf · doi:10.1016/j.ins.2021.04.066

published in Information Sciences 570, 298-305 (Elsevier BV) · Accepted for publication at Elsevier - Information Sciences

openalex publication_date 2021/04/20 · arxiv created 2021/05/01 · arxiv updated 2021/05/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06

Abstract

Feature ranking and selection is a widely used approach in various applications of supervised dimensionality reduction in discriminative machine learning. Nevertheless there exists significant evidence on feature ranking and selection algorithms based on any criterion leading to potentially sub-optimal solutions for class separability. In that regard, we introduce emerging information theoretic feature transformation protocols as an end-to-end neural network training approach. We present a dimensionality reduction network (MMINet) training procedure based on the stochastic estimate of the mutual information gradient. The network projects high-dimensional features onto an output feature space where lower dimensional representations of features carry maximum mutual information with their associated class labels. Furthermore, we formulate the training objective to be estimated non-parametrically with no distributional assumptions. We experimentally evaluate our method with applications to high-dimensional biological data sets, and relate it to conventional feature selection algorithms to form a special case of our approach.

Citations