vix.ing · top · new · best · stats

Efficient Statistics, in High Dimensions, from Truncated Samples

2018/09/11 by Constantinos Daskalakis, Daskalakis, Constantinos, Themis Gouleakis +5 · 12 citations
Computer Science · Mathematics · #Computation (stat.CO) #Data Structures and Algorithms (cs.DS) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Statistics Theory (math.ST) #cs.DS #cs.LG #math.ST #stat.CO #stat.ML #stat.TH

paper · pdf · doi:10.48550/arxiv.1809.03986

Appeared at 59th Annual IEEE Symposium on Foundations of Computer Science (FOCS), 2018

arxiv created 2020/10/22 · arxiv updated 2020/10/26

Abstract

We provide an efficient algorithm for the classical problem, going back to Galton, Pearson, and Fisher, of estimating, with arbitrary accuracy the parameters of a multivariate normal distribution from truncated samples. Truncated samples from a d-variate normal \cal N(\mathbfμ,\mathbfΣ) means a samples is only revealed if it falls in some subset S ⊆ ℝd; otherwise the samples are hidden and their count in proportion to the revealed samples is also hidden. We show that the mean \mathbfμ and covariance matrix \mathbfΣ can be estimated with arbitrary accuracy in polynomial-time, as long as we have oracle access to S, and S has non-trivial measure under the unknown d-variate normal distribution. Additionally we show that without oracle access to S, any non-trivial estimation is impossible.

Cited by

Related