vix.ing · top · new · best · stats · spec

Generalization error bounds in semi-supervised classification under the cluster assumption

2006/04/11 by Philippe Rigollet, Rigollet, Philippe · 1 citation
Computer Science · Mathematics · #Advanced Statistical Methods and Models #Bayesian Methods and Mixture Models #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning and Data Classification #Statistics Theory (math.ST)

paper · doi:10.48550/arxiv.math/0604233

openalex publication_date 2006/04/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We consider semi-supervised classification when part of the available data is unlabeled. These unlabeled data can be useful for the classification problem when we make an assumption relating the behavior of the regression function to that of the marginal distribution. Seeger (2000) proposed the well-known "cluster assumption" as a reasonable one. We propose a mathematical formulation of this assumption and a method based on density level sets estimation that takes advantage of it to achieve fast rates of convergence both in the number of unlabeled examples and the number of labeled examples.

Citations

Cited by

Related