2024/08/24 by Buxin Su, Jiayao Zhang, Su, Buxin +15 · 1 voice · 5 citations
Computer Science · Decision Sciences · Mathematics · Medicine · Psychology · Social Sciences · #Applications (stat.AP) #Computational and Text Analysis Methods #Computer Science and Game Theory (cs.GT) #Computer science #Digital Libraries (cs.DL) #FOS: Computer and information sciences #Information retrieval #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mathematics education #Medical education #Medicine #Natural Language Processing Techniques #Psychology #Ranking (information retrieval) #Reliability and Agreement in Measurement #cs.DL #cs.GT #cs.LG #stat.AP #stat.ML
paper · pdf · doi:10.48550/arxiv.2408.13430
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/08/24 · arxiv published 2024/08/24 · openalex created_date 2024/09/21 · arxiv updated 2025/09/23 · openalex updated_date 2026/08/06
We conducted an experiment during the review process of the 2023 International Conference on Machine Learning (ICML), asking authors with multiple submissions to rank their papers based on perceived quality. In total, we received 1,342 rankings, each from a different author, covering 2,592 submissions. In this paper, we present an empirical analysis of how author-provided rankings could be leveraged to improve peer review processes at machine learning conferences. We focus on the Isotonic Mechanism, which calibrates raw review scores using the author-provided rankings. Our analysis shows that these ranking-calibrated scores outperform the raw review scores in estimating the ground truth ``expected review scores'' in terms of both squared and absolute error metrics. Furthermore, we propose several cautious, low-risk applications of the Isotonic Mechanism and author-provided rankings in peer review, including supporting senior area chairs in overseeing area chairs' recommendations, assisting in the selection of paper awards, and guiding the recruitment of emergency reviewers.