2019/11/08 by Juri Opitz, Opitz, Juri, Sebastian Burst +1 · 13 citations
Computer Science · #FOS: Computer and information sciences #Imbalanced Data Classification Techniques #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Spam and Phishing Detection #Text and Document Classification Technologies
paper · pdf · doi:10.48550/arxiv.1911.03347
openalex publication_date 2019/11/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The 'macro F1' metric is frequently used to evaluate binary, multi-class and multi-label classification problems. Yet, we find that there exist two different formulas to calculate this quantity. In this note, we show that only under rare circumstances the two computations can be considered equivalent. More specifically, one formula well 'rewards' classifiers which produce a skewed error type distribution. In fact, the difference in outcome of the two computations can be as high as 0.5. The two computations may not only diverge in their scalar result but can also lead to different classifier rankings.