vix.ing · top · new · best · stats · spec

Risk bounds for PU learning under Selected At Random assumption

2022/01/17 by Olivier Coudray, Christine Keribin, Coudray, Olivier +5
Mathematics · #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (stat.ML) #Statistics Theory (math.ST) #math.ST #stat.ML #stat.TH

paper · pdf · doi:10.48550/arxiv.2201.06277

arxiv created 2022/01/17 · arxiv updated 2022/01/19

Abstract

Positive-unlabeled learning (PU learning) is known as a special case of semi-supervised binary classification where only a fraction of positive examples are labeled. The challenge is then to find the correct classifier despite this lack of information. Recently, new methodologies have been introduced to address the case where the probability of being labeled may depend on the covariates. In this paper, we are interested in establishing risk bounds for PU learning under this general assumption. In addition, we quantify the impact of label noise on PU learning compared to standard classification setting. Finally, we provide a lower bound on minimax risk proving that the upper bound is almost optimal.

Related