2018/03/24 by Kahkashan Afrin, Afrin, Kahkashan, Gurudev Illangovan +5
Computer Science · Health Professions · Social Sciences · #Artificial Intelligence in Healthcare #FOS: Computer and information sciences #Insurance, Mortality, Demography, Risk Management #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning in Healthcare #Medical Coding and Health Information
paper · pdf · doi:10.48550/arxiv.1803.09177
openalex publication_date 2018/03/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Accuracies of survival models for life expectancy prediction as well as\ncritical-care applications are significantly compromised due to the sparsity of\nsamples and extreme imbalance between the survival (usually, the majority) and\nmortality class sizes. While a recent random survival forest (RSF) model\novercomes the limitations of the proportional hazard assumption, an imbalance\nin the data results in an underestimation (overestimation) of the hazard of the\nmortality (survival) classes. A balanced random survival forests (BRSF) model,\nbased on training the RSF model with data generated from a synthetic minority\nsampling scheme is presented to address this gap. Theoretical results on the\neffect of balancing on prediction accuracies in BRSF are reported. Benchmarking\nstudies were conducted using five datasets with different levels of class\nimbalance from public repositories and an imbalanced dataset of 267 acute\ncardiac patients, collected at the Heart, Artery, and Vein Center of Fresno,\nCA. Investigations suggest that BRSF provides an improved discriminatory\nstrength between the survival and the mortality classes. It outperformed both\noptimized Cox (without and with balancing) and RSF with an average reduction of\n55 % in the prediction error over the next best alternative.\n