2021/06/11 by Sebastin Santy, Prasanta Bhattacharya, Santy, Sebastin +1 · 1 citation
Chemistry · Computer Science · #Artificial intelligence #Chemistry #Computation and Language (cs.CL) #Computer science #Epistemology #FOS: Computer and information sciences #Focus (optics) #Machine translation #Mechanism (biology) #Natural Language Processing Techniques #Natural language processing #Philosophy #Software Engineering Research #Topic Modeling #Translation (biology) #cs.CL
paper · pdf · doi:10.48550/arxiv.2106.06292
published in arXiv (Cornell University) (Cornell University) · pre-print
openalex publication_date 2021/06/11 · arxiv created 2022/12/30 · arxiv updated 2023/01/02 · openalex created_date 2023/01/06 · openalex updated_date 2026/07/28
Recent advances in AI and ML applications have benefited from rapid progress in NLP research. Leaderboards have emerged as a popular mechanism to track and accelerate progress in NLP through competitive model development. While this has increased interest and participation, the over-reliance on single, and accuracy-based metrics have shifted focus from other important metrics that might be equally pertinent to consider in real-world contexts. In this paper, we offer a preliminary discussion of the risks associated with focusing exclusively on accuracy metrics and draw on recent discussions to highlight prescriptive suggestions on how to develop more practical and effective leaderboards that can better reflect the real-world utility of models.