2021/03/30 by Daniel J. Wu, Wu, Daniel J., Avoy Datta +1 · 2 citations
Computer Science · Mathematics · #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Artificial intelligence #Computer science #Discretization #FOS: Computer and information sciences #Function (biology) #I.2.m #Imbalanced Data Classification Techniques #Machine Learning (cs.LG) #Machine Learning and Data Classification #Machine learning #Mathematics #Regression #Sample (material) #Simple (philosophy) #Statistics #Weight function #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2103.16591
published in arXiv (Cornell University) (Cornell University) · 4 pages, 2 figures, presented at the S2D-OLAD workshop at ICLR 2021
arxiv created 2021/03/30 · openalex publication_date 2021/03/30 · arxiv updated 2021/04/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We propose a simple method by which to choose sample weights for problems with highly imbalanced or skewed traits. Rather than naively discretizing regression labels to find binned weights, we take a more principled approach -- we derive sample weights from the transfer function between an estimated source and specified target distributions. Our method outperforms both unweighted and discretely-weighted models on both regression and classification tasks. We also open-source our implementation of this method (https://github.com/Daniel-Wu/Continuous-Weight-Balancing) to the scientific community.