vix.ing · top · new · best · stats

What the F-measure doesn't measure: Features, Flaws, Fallacies and Fixes

2015/03/22 by David Powers, David M. W. Powers, Powers, David M. W. · 2 citations
Computer Science · Mathematics · #68Q32 #68T05 #91E45 #Computation (stat.CO) #Computation and Language (cs.CL) #D.2.8 #FOS: Computer and information sciences #I.2.6 #I.2.7 #I.4.6 #I.5.1 #Information Retrieval (cs.IR) #Information Retrieval and Search Behavior #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mathematics, Computing, and Information Processing #Natural Language Processing Techniques #Neural and Evolutionary Computing (cs.NE) #acm:68Q32 #acm:68T05 #acm:91E45 #cs.CL #cs.IR #cs.LG #cs.NE #msc:68Q32 #msc:68T05 #msc:91E45 #stat.CO #stat.ML

paper · pdf · doi:10.48550/arxiv.1503.06410

19 pages

openalex publication_date 2015/03/22 · arxiv created 2019/09/12 · arxiv updated 2019/09/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The F-measure or F-score is one of the most commonly used single number measures in Information Retrieval, Natural Language Processing and Machine Learning, but it is based on a mistake, and the flawed assumptions render it unsuitable for use in most contexts! Fortunately, there are better alternatives.

Cited by

Related