Comparison of Foundation and Supervised Learning-Based Models for Detection of Referable Glaucoma from Fundus Photographs
Bolo, K.; Nguyen, T. H.; Iyengar, S.; Li, Z.; Nguyen, V.; Wong, B.; Do, J.; Ambite, J.-L.; Kesselman, C.; Daskivich, L.; Xu, B.
Show abstract
PurposeTo compare the performance of a foundation model and a supervised learning-based model for detecting referable glaucoma from fundus photographs. DesignEvaluation of diagnostic technology. Participants6,116 participants from the Los Angeles County Department of Health Services Teleretinal Screening Program. MethodsFundus photographs were labeled for referable glaucoma (cup-to-disc ratio [≥] 0.6) by certified optometrists. Four deep learning models were trained on cropped and uncropped images (Training N = 8,996; Validation N = 3,002) using two architectures: a vision transformer with self-supervised pretraining on fundus photographs (RETFound) and a convolutional neural network (VGG-19). Models were evaluated on a held-out test set (N = 1,000) labeled by glaucoma specialists and an external test set (N = 300) from University of Southern California clinics. Performance was assessed while varying training set size and stratifying by demographic factors. xRAI was used for saliency mapping. Main Outcome MeasuresArea under the receiver operating characteristic curve (AUC-ROC) and threshold-specific metrics. ResultsThe cropped image VGG-19 model achieved the highest AUC-ROC (0.924 [0.907-0.940]), which was comparable (p = 0.07) to the cropped image RETFound model (0.911 [0.892-0.930]), which achieved the highest Youden-optimal performance (sensitivity 82.6%, specificity 88.2%) and F1 score (0.801). Cropped image models outperformed their uncropped counterparts within each architecture (p < 0.001 for AUC-ROC comparisons). RETFound models had a performance advantage when trained on smaller datasets (N < 2000 images), and the uncropped image RETFound model performed best on external data (p < 0.001 for AUC-ROC comparisons). The cropped image RETFound model performed consistently across ethnic groups (p = 0.20), while the others did not (p < 0.04); performance did not vary by age or gender. Saliency maps for both architectures consistently included the optic nerve. ConclusionWhile both RETFound and VGG-19 models performed well for classification of referable glaucoma, foundation models may be preferable when training data is limited and when domain shift is expected. Training models using images cropped to the region of the optic nerve improves performance regardless of architecture but may reduce model generalizability.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Relating Standardized Automated Perimetry Performed with Stimulus Sizes III and V in Eyes With Field Loss due to Glaucoma and NAION 96%
- Rates of Glaucoma Progression Derived from Linear Mixed Models Using Varied Random Effect Distributions 96%
- Reliability of retinal pathology quantification in age-related macular degeneration: Implications for clinical trials and machine learning applications 96%
Similar papers in this journal
Similar papers in this journal
- Automated Expert-level Scleral Spur Detection and Quantitative Biometric Analysis on the ANTERION Anterior Segment OCT System 97%
- Macula structural and vascular differences in glaucoma eyes with and without high axial myopia 96%
- Unveiling the Clinical Incapabilities: A Benchmarking Study of GPT-4V(ision) for Ophthalmic Multimodal Image Analysis 95%
Similar papers in this journal
- Self-supervised contrastive learning improves machine learning discrimination of full thickness macular holes from epiretinal membranes in retinal OCT scans 96%
- An Inherently Interpretable AI model improves Screening Speed and Accuracy for Early Diabetic Retinopathy 95%
- Detecting papilloedema as a marker of raised intracranial pressure using artificial intelligence: a systematic review 95%
Similar papers in this journal
- Glaucoma Detection and Staging from Visual Field Images using Machine Learning Techniques 96%
- Prediction of the ectasia screening index from raw Casia2 volume data for keratoconus identification by using convolutional neural networks 95%
- Towards implementation of AI in New Zealand national screening program: Cloud-based, Robust, and Bespoke 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.