vix.ing · top · new · best · stats · spec

Position: Stop Chasing the C-index when Evaluating Survival Analysis Models

2025/06/02 by Christian Marius Lillelund, Shi-ang Qi, Lillelund, Christian Marius +5 · 1 voice · 2 citations
Computer Science · Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare #Methodology (stat.ME) #cs.LG #stat.ME

paper · pdf · doi:10.48550/arxiv.2506.02075

openalex publication_date 2025/06/02 · arxiv published 2025/06/02 · openalex created_date 2025/10/14 · arxiv updated 2026/05/31 · openalex updated_date 2026/07/28

Abstract

The current state of evaluation in survival analysis is plagued by the persistent use of evaluation metrics in ways that are misaligned with the stated modeling objective. In addition, many such evaluations are based on censoring assumptions that are left implicit or unjustified. This means that the reported performance can be misleading and may fail to answer the scientific or modeling question the evaluation was intended to address. In this position paper, we critically examine evaluation practices in survival analysis and highlight how censoring makes evaluation fundamentally different from standard regression or classification. We place particular focus on concordance-based measures, such as the C-index, which we show are heavily overused in the literature. To help identify appropriate metrics, we propose a set of key desiderata and introduce a double-helix ladder, in which valid evaluation requires alignment between metric and modeling assumptions. Through controlled experiments, we show that violations of this alignment can lead to misleading model comparisons. We conclude by providing practical guidance on how to evaluate a survival model.

Cited by

Discussions

Related