2020/04/29 by Arjun Desai, Desai, Arjun D., Francesco Calivá +54
Medicine · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Osteoarthritis Treatment and Mechanisms #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2004.14003
openalex publication_date 2020/04/29 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
Purpose: To organize a knee MRI segmentation challenge for characterizing the\nsemantic and clinical efficacy of automatic segmentation methods relevant for\nmonitoring osteoarthritis progression.\n Methods: A dataset partition consisting of 3D knee MRI from 88 subjects at\ntwo timepoints with ground-truth articular (femoral, tibial, patellar)\ncartilage and meniscus segmentations was standardized. Challenge submissions\nand a majority-vote ensemble were evaluated using Dice score, average symmetric\nsurface distance, volumetric overlap error, and coefficient of variation on a\nhold-out test set. Similarities in network segmentations were evaluated using\npairwise Dice correlations. Articular cartilage thickness was computed per-scan\nand longitudinally. Correlation between thickness error and segmentation\nmetrics was measured using Pearson's coefficient. Two empirical upper bounds\nfor ensemble performance were computed using combinations of model outputs that\nconsolidated true positives and true negatives.\n Results: Six teams (T1-T6) submitted entries for the challenge. No\nsignificant differences were observed across all segmentation metrics for all\ntissues (p=1.0) among the four top-performing networks (T2, T3, T4, T6). Dice\ncorrelations between network pairs were high (>0.85). Per-scan thickness errors\nwere negligible among T1-T4 (p=0.99) and longitudinal changes showed minimal\nbias (<0.03mm). Low correlations (<0.41) were observed between segmentation\nmetrics and thickness error. The majority-vote ensemble was comparable to top\nperforming networks (p=1.0). Empirical upper bound performances were similar\nfor both combinations (p=1.0).\n Conclusion: Diverse networks learned to segment the knee similarly where high\nsegmentation accuracy did not correlate to cartilage thickness accuracy. Voting\nensembles did not outperform individual networks but may help regularize\nindividual models.\n