2019/09/21 by Luma Omar, Omar, Luma, Ioannis Ivrissimtzis +1
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Data Mining Algorithms and Applications #FOS: Computer and information sciences #Imbalanced Data Classification Techniques #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification
paper · pdf · doi:10.48550/arxiv.1909.09816
openalex publication_date 2019/09/21 · openalex created_date 2022/07/19 · openalex updated_date 2026/07/28
Most binary classifiers work by processing the input to produce a scalar\nresponse and comparing it to a threshold value. The various measures of\nclassifier performance assume, explicitly or implicitly, probability\ndistributions Ps and Pn of the response belonging to either class,\nprobability distributions for the cost of each type of misclassification, and\ncompute a performance score from the expected cost.\n In machine learning, classifier responses are obtained experimentally and\nperformance scores are computed directly from them, without any assumptions on\nPs and Pn. Here, we argue that the omitted step of estimating theoretical\ndistributions for Ps and Pn can be useful. In a biometric security\nexample, we fit beta distributions to the responses of two classifiers, one\nbased on logistic regression and one on ANNs, and use them to establish a\ncategorisation into a small number of classes with different extremal\nbehaviours at the ends of the ROC curves.\n