2021/04/27 by Wenwen Min, Min, Wenwen, Taosheng Xu +5
Biochemistry, Genetics and Molecular Biology · Computer Science · Engineering · #Blind Source Separation Techniques #FOS: Computer and information sciences #Gene expression and cancer classification #Machine Learning (cs.LG) #Sparse and Compressive Sensing Techniques
paper · pdf · doi:10.48550/arxiv.2104.13171
openalex publication_date 2021/04/27 · openalex created_date 2021/05/10 · openalex updated_date 2026/07/28
Non-negative matrix factorization (NMF) is a powerful tool for dimensionality reduction and clustering. Unfortunately, the interpretation of the clustering results from NMF is difficult, especially for the high-dimensional biological data without effective feature selection. In this paper, we first introduce a row-sparse NMF with ℓ2,0-norm constraint (NMF_ℓ20), where the basis matrix W is constrained by the ℓ2,0-norm, such that W has a row-sparsity pattern with feature selection. It is a challenge to solve the model, because the ℓ2,0-norm is non-convex and non-smooth. Fortunately, we prove that the ℓ2,0-norm satisfies the Kurdyka-Łojasiewicz property. Based on the finding, we present a proximal alternating linearized minimization algorithm and its monotone accelerated version to solve the NMF_ℓ20 model. In addition, we also present a orthogonal NMF with ℓ2,0-norm constraint (ONMF_ℓ20) to enhance the clustering performance by using a non-negative orthogonal constraint. We propose an efficient algorithm to solve ONMF_ℓ20 by transforming it into a series of constrained and penalized matrix factorization problems. The results on numerical and scRNA-seq datasets demonstrate the efficiency of our methods in comparison with existing methods.