2019/05/14 by Sushrut Karmalkar, Karmalkar, Sushrut, Adam R. Klivans +3 · 5 citations
Computer Science · Engineering · Mathematics · #Data Structures and Algorithms (cs.DS) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Sparse and Compressive Sensing Techniques #cs.DS #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1905.05679
28 Pages
openalex publication_date 2019/05/14 · arxiv created 2019/05/30 · arxiv updated 2019/05/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We give the first polynomial-time algorithm for robust regression in the list-decodable setting where an adversary can corrupt a greater than 1/2 fraction of examples. For any α< 1, our algorithm takes as input a sample \(xi,yi)\i ≤ n of n linear equations where αn of the equations satisfy yi = ⟨ xi,ℓ^*⟩ +ζ for some small noise ζ and (1-α)n of the equations are \em arbitrarily chosen. It outputs a list L of size O(1/α) - a fixed constant - that contains an ℓ that is close to ℓ^*. Our algorithm succeeds whenever the inliers are chosen from a certifiably anti-concentrated distribution D. In particular, this gives a (d/α)O(1/α8) time algorithm to find a O(1/α) size list when the inlier distribution is standard Gaussian. For discrete product distributions that are anti-concentrated only in regular directions, we give an algorithm that achieves similar guarantee under the promise that ℓ^* has all coordinates of the same magnitude. To complement our result, we prove that the anti-concentration assumption on the inliers is information-theoretically necessary. Our algorithm is based on a new framework for list-decodable learning that strengthens the `identifiability to algorithms' paradigm based on the sum-of-squares method. In an independent and concurrent work, Raghavendra and Yau also used the Sum-of-Squares method to give a similar result for list-decodable regression.