vix.ing · top · new · best · stats · spec

Telling Two Distributions Apart: a Tight Characterization

2011/10/14 by Eyal Even-Dar, Dar, Eyal Even, M. Sandler +1
Computer Science · #Data Stream Mining Techniques #Data Structures and Algorithms (cs.DS) #FOS: Computer and information sciences #Machine Learning and Algorithms #Machine Learning and Data Classification

paper · pdf · doi:10.48550/arxiv.1110.3100

openalex publication_date 2011/10/14 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28

Abstract

We consider the problem of distinguishing between two arbitrary black-box distributions defined over the domain [n], given access to s samples from both. It is known that in the worst case O(n2/3) samples is both necessary and sufficient, provided that the distributions have L1 difference of at least ε. However, it is also known that in many cases fewer samples suffice. We identify a new parameter, that provides an upper bound on how many samples needed, and present an efficient algorithm that requires the number of samples independent of the domain size. Also for a large subclass of distributions we provide a lower bound, that matches our upper bound up to a poly-logarithmic factor.

Related