2021/11/11 by Shun Xu, Ming Yuan, Xu, Shun +1
Computer Science · Engineering · #Blind Source Separation Techniques #FOS: Computer and information sciences #Image and Signal Denoising Methods #Machine Learning (stat.ML) #Methodology (stat.ME) #Sparse and Compressive Sensing Techniques
paper · pdf · doi:10.48550/arxiv.2111.06302
openalex publication_date 2021/11/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this note, we investigate how well we can reconstruct the best rank-r approximation of a large matrix from a small number of its entries. We show that even if a data matrix is of full rank and cannot be approximated well by a low-rank matrix, its best low-rank approximations may still be reliably computed or estimated from a small number of its entries. This is especially relevant from a statistical viewpoint: the best low-rank approximations to a data matrix are often of more interest than itself because they capture the more stable and oftentimes more reproducible properties of an otherwise complicated data-generating model. In particular, we investigate two agnostic approaches: the first is based on spectral truncation; and the second is a projected gradient descent based optimization procedure. We argue that, while the first approach is intuitive and reasonably effective, the latter has far superior performance in general. We show that the error depends on how close the matrix is to being of low rank. Both theoretical and numerical evidence is presented to demonstrate the effectiveness of the proposed approaches.