2024/02/12 by Anindya De, Huan Li, De, Anindya +5 · 2 citations
Computer Science · Mathematics · #Complexity and Algorithms in Graphs #Computational Complexity (cs.CC) #Data Structures and Algorithms (cs.DS) #FOS: Computer and information sciences #Machine Learning and Algorithms #Markov Chains and Monte Carlo Methods
paper · pdf · doi:10.48550/arxiv.2402.08133
openalex publication_date 2024/02/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We consider the following basic, and very broad, statistical problem: Given a known high-dimensional distribution \cal D over ℝn and a collection of data points in ℝn, distinguish between the two possibilities that (i) the data was drawn from \cal D, versus (ii) the data was drawn from \cal D|S, i.e. from \cal D subject to truncation by an unknown truncation set S ⊆ ℝn. We study this problem in the setting where \cal D is a high-dimensional i.i.d. product distribution and S is an unknown degree-d polynomial threshold function (one of the most well-studied types of Boolean-valued function over ℝn). Our main results are an efficient algorithm when \cal D is a hypercontractive distribution, and a matching lower bound: \bullet For any constant d, we give a polynomial-time algorithm which successfully distinguishes \cal D from \cal D|S using O(nd/2) samples (subject to mild technical conditions on \cal D and S); \bullet Even for the simplest case of \cal D being the uniform distribution over \+1, -1\n, we show that for any constant d, any distinguishing algorithm for degree-d polynomial threshold functions must use Ω(nd/2) samples.