2016/10/11 by Patrice Bertail, Stephan, Clémençon, Bertail, Patrice +2 · 1 citation
Mathematics · Computer Science · #Statistical Methods and Inference #Advanced Statistical Methods and Models #Machine Learning and Algorithms
paper · pdf · doi:10.48550/arxiv.1610.03316
The generalization ability of minimizers of the empirical risk in the context\nof binary classification has been investigated under a wide variety of\ncomplexity assumptions for the collection of classifiers over which\noptimization is performed. In contrast, the vast majority of the works\ndedicated to this issue stipulate that the training dataset used to compute the\nempirical risk functional is composed of i.i.d. observations. Beyond the cases\nwhere training data are drawn uniformly without replacement among a large\ni.i.d. sample or modelled as a realization of a weakly dependent sequence of\nr.v.'s, statistical guarantees when the data used to train a classifier are\ndrawn by means of a more general sampling/survey scheme and exhibit a complex\ndependence structure have not been documented yet. It is the main purpose of\nthis paper to show that the theory of empirical risk minimization can be\nextended to situations where statistical learning is based on survey samples\nand knowledge of the related inclusion probabilities. Precisely, we prove that\nminimizing a weighted version of the empirical risk, refered to as the\nHorvitz-Thompson risk (HT risk), over a class of controlled complexity lead to\na rate for the excess risk of the order O\ℙ((\κN (\log\nN)/n)1/2) with \κN=(n/N)/\mini\≤ N\πi, when data are sampled\nby means of a rejective scheme of (deterministic) size n within a statistical\npopulation of cardinality N\≥ n, a generalization of basic it sampling\nwithout replacement with unequal probability weights \πi>0. Extension to\nother sampling schemes are then established by a coupling argument. Beyond\ntheoretical results, numerical experiments are displayed in order to show the\nrelevance of HT risk minimization and that ignoring the sampling scheme used to\ngenerate the training dataset may completely jeopardize the learning procedure.\n