2026/07/23 by Szilárd Nemes
Mathematics · #Statistical Methods and Inference #Advanced Causal Inference Techniques #Statistical Methods and Bayesian Inference
paper · doi:10.1080/00031305.2026.2701342
The Kaplan–Meier estimator is standard for censored survival data, but information lost to censoring is rarely made explicit. The practical question is not only how many subjects remain at risk, but also how much information remains for inference. We develop a variance-decomposition framework that makes this loss explicit by separating complete-follow-up variability from the additional component induced by censoring. This yields pointwise measures of information retention: a variance-inflation factor, relative efficiency, and an effective sample size, ESŜ(t). ESŜ(t) has a binomial-equivalent interpretation: it is the size of a hypothetical complete-follow-up binomial experiment with survival probability ŜKM(t) whose variance equals the Greenwood variance. The framework also yields a pointwise censoring-component interval, showing where the survival estimate from the same sample would plausibly have fallen if censoring had not occurred. Simulations show that ESŜ(t) is well calibrated and that the censoring-component interval has near-nominal pointwise coverage. Examples from clinical trials illustrate that risk-set size and ESŜ(t) answer different questions: a small number at risk does not imply low precision when risk-set depletion is driven primarily by observed events rather than censoring. The proposed tools provide practical data-maturity summaries for judging how much information remains over time and how reliable later Kaplan–Meier inference is.