why choose us

300×250 Ad Slot

Research Article: Explainability-informed patient-disjoint evaluation of deep learning for plus disease in retinopathy of prematurity with an ophthalmologist reader study

Date Published: 2026-09-30

Abstract:
Plus disease is a treatment-relevant sign of retinopathy of prematurity, but fundus images from the same infant may differ in image quality, posterior-pole coverage, and visible vascular severity. We evaluated deep-learning classification using a patient-disjoint design, within-infant variability, and an ophthalmologist reader study. The primary analysis included 732 Phoenix ICON fundus images from 38 infants, assigned exclusively to a 634-image development pool or a 98-image hold-out set from eight previously unseen infants. Five pretrained convolutional neural networks were evaluated at the image level. Hold-out predictions were grouped by infant and laterality to characterize descriptive within-infant variation. A historical secondary analysis used a random image-level partition without enforced patient separation and was considered exploratory. Two ophthalmologists independently graded 30 unique isolated images for inter-observer analysis; 10 were presented twice, resulting in 40 presentations for intra-observer repeatability. DenseNet121 achieved an image-level ROC-AUC of 0.929; however, infant-equal-weighted image-level ROC-AUC was 0.905 and leave-one-infant-out ROC-AUC ranged from 0.838 to 0.989, demonstrating sensitivity to the small clustered hold-out cohort. At a development-defined threshold of 0.0184, sensitivity increased to 98.0%, while specificity decreased to 46.8%. Within-infant probability variability was observed across all five CNN architectures. Ophthalmologists 1 and 2 assessed 14/18 and 17/18 plus-reference images, respectively, as Healthy. Exact three-category inter-observer agreement was 26/30 (86.7%), and both reproduced the same severity grade for all 10 repeated images. Patient-disjoint evaluation revealed substantial uncertainty in image-level performance and variability among images from the same infant. Cluster-aware analyses showed dependence on the small hold-out cohort, while the development-defined operating point demonstrated a marked sensitivity–specificity trade-off. The reader study showed that an examination-derived reference label and the diagnostic evidence visible in an isolated image are not equivalent, despite high inter-observer agreement and complete categorical repeatability. Because examination identifiers were unavailable, within-infant analyses were interpreted descriptively and not as infant or examination-level diagnostic performance. Clinically useful systems should assess image quality, select informative views, and integrate evidence across multiple images and both eyes. External prospective multicentre validation is required before clinical implementation.

Introduction:
Plus disease is a treatment-relevant sign of retinopathy of prematurity, but fundus images from the same infant may differ in image quality, posterior-pole coverage, and visible vascular severity. We evaluated deep-learning classification using a patient-disjoint design, within-infant variability, and an ophthalmologist reader study.

Read more

300×250 Ad Slot