Development and validation of a lesion-supervised deep learning system for diabetic retinopathy grading according to UK national screening criteria
Chowdhury, P. N.; Akter, Y.; Chowdhury, P.; Kaur, A.; Uddin, M.; Chowdhury, A.; Chowdhury, P. K.; Muqit, M.
Show abstract
BackgroundDiabetic retinopathy (DR) is the leading cause of preventable blindness among working-age adults worldwide, yet screening coverage remains inadequate, particularly in low-and middle-income countries. Automated deep learning systems offer potential to address the global shortage of expert graders, but most existing models lack lesion-level interpretability and are not aligned with established clinical referral frameworks. We developed and validated DRAGS (Diabetic Retinopathy Automated Grading System), a hybrid deep learning model that grades DR according to the UK Diabetic Eye Screening Programme (DESP) classification and provides lesion-level explainability. MethodsWe trained and validated a DenseNet-201-based convolutional neural network on 20,281 anonymised fundus images from two tertiary eye care institutions in Bangladesh. Images were graded by fellowship-trained retinal specialists using the UK DESP framework, resulting in 10 clinically interpretable classes that combine retinopathy grade (R0-R3) and maculopathy status (M0/M1). A companion dataset of 2,936 pixel-level lesion masks spanning nine pathological categories was used to train a parallel multi-label lesion-detection head. The dataset was partitioned 70:15:15 (patient-stratified). Performance was evaluated using macro-averaged AUROC (DeLong estimator), sensitivity, specificity, F1 score, quadratically weighted Cohens {kappa}, and expected calibration error (ECE), with 95% CIs from 2000 bootstrap resamples. Grad-CAM spatial alignment with ground-truth lesion masks was assessed using Dice and IoU. This study follows the TRIPOD+AI reporting guidelines. FindingsOn the held-out test set (Component I: n = 3,044; Component II: n {approx} 440), DRAGS achieved class-wise precision, recall, and F1 scores ranging from 0{middle dot}88 to 0{middle dot}99 across all ten UK DESP grades, with advanced proliferative stages (R3-M0, R3-M1) consistently exceeding 0{middle dot}95. Overall accuracy was approximately 91{middle dot}1% and quadratically weighted Cohens {kappa} was approximately 0{middle dot}90. For referable versus non-referable DR, sensitivity was 90{middle dot}7% and specificity was 91{middle dot}9%. The companion lesion-detection head achieved macro-averaged sensitivity of 93{middle dot}9%, specificity of 99{middle dot}5%, and AUC of 0{middle dot}997 across nine lesion classes; seven of nine classes achieved AUC = 1{middle dot}00. Grad-CAM activations showed progressive spatial shift from diffuse (normal) to lesion-dense peripheral patterns (proliferative DR), with maximal agreement for microaneurysms and exudates. Mean inference time was 110-160 ms per image. InterpretationDRAGS demonstrates high diagnostic accuracy for nine-class UK DESP-aligned DR grading, with clinically interpretable lesion-level explainability on a large real-world LMIC dataset. External validation and prospective clinical evaluation are warranted before deployment. FundingThe present study received no funding.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AutoMorph: Automated Retinal Vascular Morphology Quantification via a Deep Learning Pipeline 96%
- Current applications of artificial intelligence for Fuchs endothelial corneal dystrophy: a systematic review 94%
- Reliability of retinal pathology quantification in age-related macular degeneration: Implications for clinical trials and machine learning applications 94%
Similar papers in this journal
- An Inherently Interpretable AI model improves Screening Speed and Accuracy for Early Diabetic Retinopathy 97%
- Self-supervised contrastive learning improves machine learning discrimination of full thickness macular holes from epiretinal membranes in retinal OCT scans 96%
- Detecting papilloedema as a marker of raised intracranial pressure using artificial intelligence: a systematic review 93%
Similar papers in this journal
- Annotation-free multi-organ anomaly detection in abdominal CT using free-text radiology reports: A multi-center retrospective study 90%
- Artificial Intelligence-Enhanced Comprehensive Assessment of the Aortic Valve Stenosis Continuum in Echocardiography 90%
- Deep Learning Prediction of Biomarkers from Echocardiogram Videos 89%
Similar papers in this journal
- A user-friendly tool for cloud-based whole slide image segmentation, with examples from renal histopathology 92%
- LUNAR: A Deep Learning Model to Predict Glioma Recurrence Using Integrated Genomic and Clinical Data 89%
- Harnessing Deep Learning to Detect Bronchiolitis Obliterans Syndrome from Chest CT 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.