vix.ing · top · new · best · stats

The Real-World-Weight Cross-Entropy Loss Function: Modeling the Costs of Mislabeling

2019/12/27 by Yaoshiang Ho, Samuel Wookey · 827 citations
Computer Science · Mathematics · #Anomaly Detection Techniques and Applications #Artificial intelligence #Binary classification #Categorical variable #Computer science #Data mining #Deep learning #Entropy (arrow of time) #Function (biology) #Imbalanced Data Classification Techniques #MNIST database #Machine Learning and Data Classification #Machine learning #Metric (unit) #Pattern recognition (psychology) #Probabilistic logic #Support vector machine #cs.AI #cs.LG #stat.ML

paper · pdf · doi:10.1109/access.2019.2962617

published in IEEE Access 8, 4806-4813 (Institute of Electrical and Electronics Engineers) · Submitted to IEEE Access

openalex publication_date 2019/12/27 · arxiv created 2020/01/03 · arxiv updated 2020/01/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06

Abstract

In this paper, we propose a new metric to measure goodness-of-fit for classifiers: the Real World Cost function. This metric factors in information about a real world problem, such as financial impact, that other measures like accuracy or F1 do not. This metric is also more directly interpretable for users. To optimize for this metric, we introduce the Real-World-Weight Cross-Entropy loss function, in both binary classification and single-label multiclass classification variants. Both variants allow direct input of real world costs as weights. For single-label, multiclass classification, our loss function also allows direct penalization of probabilistic false positives, weighted by label, during the training of a machine learning model. We compare the design of our loss function to the binary cross-entropy and categorical cross-entropy functions, as well as their weighted variants, to discuss the potential for improvement in handling a variety of known shortcomings of machine learning, ranging from imbalanced classes to medical diagnostic error to reinforcement of social bias. We create scenarios that emulate those issues using the MNIST data set and demonstrate empirical results of our new loss function. Finally, we discuss our intuition about why this approach works and sketch a proof based on Maximum Likelihood Estimation.

Citations

Cited by