2024/03/27 by Malka Gorfine, Gorfine, Malka, David M. Zucker +3
Computer Science · Social Sciences · #Computational and Text Analysis Methods #FOS: Computer and information sciences #Machine Learning in Healthcare #Methodology (stat.ME)
paper · pdf · doi:10.48550/arxiv.2403.18464
openalex publication_date 2024/03/27 · openalex created_date 2024/03/29 · openalex updated_date 2026/07/28
Many countries have established population-based biobanks, which are being used increasingly in epidemiolgical and clinical research. These biobanks offer opportunities for large-scale studies addressing questions beyond the scope of traditional clinical trials or cohort studies. However, using biobank data poses new challenges. Typically, biobank data is collected from a study cohort recruited over a defined calendar period, with subjects entering the study at various ages falling between cL and cU. This work focuses on biobank data with individuals reporting disease-onset age upon recruitment, termed prevalent data, along with individuals initially recruited as healthy, and their disease onset observed during the follow-up period. We propose a novel cumulative incidence function (CIF) estimator that efficiently incorporates prevalent cases, in contrast to existing methods, providing two advantages: (1) increased efficiency, and (2) CIF estimation for ages before the lower limit, cL.