2017/06/25 by Vatsal Sharan, Kai Sheng Tai, Sharan, Vatsal +5 · 1 citation
Computer Science · Engineering · Mathematics · #Algorithm #Artificial Intelligence (cs.AI) #Artificial intelligence #Blind Source Separation Techniques #Combinatorics #Compressed sensing #Computer science #Domain (mathematical analysis) #FOS: Computer and information sciences #Factorization #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mathematics #Matrix (chemical analysis) #Matrix decomposition #Non-negative matrix factorization #Pattern recognition (psychology) #Rank (graph theory) #Regular polygon #Sparse and Compressive Sensing Techniques #Sparse matrix #Tensor (intrinsic definition) #Tensor decomposition and applications #Theoretical computer science #cs.AI #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1706.08146
published in arXiv (Cornell University) (Cornell University) · Updates for ICML'19 camera-ready
openalex publication_date 2017/06/25 · arxiv created 2019/05/27 · arxiv updated 2019/05/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
What learning algorithms can be run directly on compressively-sensed data? In this work, we consider the question of accurately and efficiently computing low-rank matrix or tensor factorizations given data compressed via random projections. We examine the approach of first performing factorization in the compressed domain, and then reconstructing the original high-dimensional factors from the recovered (compressed) factors. In both the matrix and tensor settings, we establish conditions under which this natural approach will provably recover the original factors. While it is well-known that random projections preserve a number of geometric properties of a dataset, our work can be viewed as showing that they can also preserve certain solutions of non-convex, NP-Hard problems like non-negative matrix factorization. We support these theoretical results with experiments on synthetic data and demonstrate the practical applicability of compressed factorization on real-world gene expression and EEG time series datasets.