2021/01/27 by Sebastian Farquhar, Yarin Gal, Farquhar, Sebastian +3 · 24 citations
Computer Science · Engineering · Mathematics · #Active learning (machine learning) #Algorithms and Data Compression #Artificial intelligence #Computer science #Engineering #FOS: Computer and information sciences #Inductive bias #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Machine Learning and Data Classification #Machine learning #Mechanism (biology) #Multi-task learning #Population #Statistical learning #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2101.11665
published in arXiv (Cornell University) (Cornell University) · Published at ICLR 2021 (Spotlight)
openalex publication_date 2021/01/27 · arxiv created 2021/05/31 · arxiv updated 2021/06/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Active learning is a powerful tool when labelling data is expensive, but it introduces a bias because the training data no longer follows the population distribution. We formalize this bias and investigate the situations in which it can be harmful and sometimes even helpful. We further introduce novel corrective weights to remove bias when doing so is beneficial. Through this, our work not only provides a useful mechanism that can improve the active learning approach, but also an explanation of the empirical successes of various existing approaches which ignore this bias. In particular, we show that this bias can be actively helpful when training overparameterized models -- like neural networks -- with relatively little data.