Automated Quantification of Decreased FAF in Stargardt Disease: Validation of a Novel Method Compared to Manual Grading Standards
Ahmed, M. I.; Yucel, H.; Afridi, R.; de Guimaraes, T. A.; Sendino-Tenorio, I.; Nguyen, N. V.; Kiran, R.; Khan, S.; Khan, U.; un Nisa, S. S.; Campigotto, M.; Hariri, A.; Michaelides, M.; Scholl, H. P.; Mata, N.; Nguyen, Q. D.; Sepah, Y. J.
Show abstract
PurposeTo evaluate the repeatability and reproducibility of a novel automated method compared with manual segmentation for measuring decreased autofluorescence (DAF) and definitely decreased autofluorescence (DDAF) in fundus autofluorescence (FAF) images of patients with Stargardt disease. DesignCross-sectional reproducibility and agreement study. ParticipantsA total of 316 eyes from 158 genetically confirmed Stargardt patients were analyzed. For intra-grader repeatability, 114 FAF images were reassessed in a masked, repeated-measures design. MethodsDAF and DDAF lesion areas were independently quantified by five certified graders using either manual delineation with Heidelberg RegionFinder or a threshold-based automated algorithm. Agreement and repeatability were assessed using intraclass correlation coefficients (ICC), standard error of measurement (SEM), minimal detectable change (MDC), Lins concordance correlation coefficient (CCC), Bland-Altman plots, and Passing-Bablok regression. Both raw and square-root-transformed lesion areas were evaluated. Main Outcome MeasuresRepeatability (intra-grader ICC, SEM, MDC), reproducibility (inter-grader ICC), and agreement (CCC, bias in regression analysis) between and within manual and automated methods. ResultsThe automated method achieved excellent intra-grader repeatability for both DAF and DDAF (ICCs [≥]0.988, SEM [≤]0.71 mm{superscript 2}, MDC [≤]1.98 mm{superscript 2}), with minimal operator influence. Manual measurements showed variable repeatability (DAF ICCs 0.909-0.974; DDAF ICCs as low as 0.837), with square-root transformation reducing SEM and MDC. Inter-grader reproducibility was highest for automated methods (ICC = 0.989-0.992), whereas manual methods ranged from 0.764-0.939 (raw) and 0.867-0.922 (transformed). Cross-method agreement was strong (CCC = 0.91-0.96), though minor proportional and constant bias was observed in raw DAF data. ConclusionsThe automated approach provides near-perfect repeatability and high agreement with manual grading, offering a scalable, objective alternative for quantifying hypo-autofluorescent lesions in Stargardt disease. Manual methods are generally reliable but more variable, especially for DDAF, and benefit from square-root transformation.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Reliability of retinal pathology quantification in age-related macular degeneration: Implications for clinical trials and machine learning applications 94%
- Visual field evaluation using Zippy Adaptive Threshold Algorithm (ZATA) Standard and ZATA Fast in patients with glaucoma and healthy individuals 94%
- Low-cost, Smartphone-based Specular Imaging and Automated Analysis of the Corneal Endothelium 94%
Similar papers in this journal
- Development of the Advised Protocol for OCT Study Terminology and Elements Anterior Segment OCT extension reporting guidelines: APOSTEL-AS 93%
- Automation improves repeatability of retinal oximetry measurements 93%
- Prediction of the ectasia screening index from raw Casia2 volume data for keratoconus identification by using convolutional neural networks 93%
Similar papers in this journal
- Automated Expert-level Scleral Spur Detection and Quantitative Biometric Analysis on the ANTERION Anterior Segment OCT System 95%
- Autonomous Screening for Laser Photocoagulation in Fundus Images Using Deep Learning 94%
- Unveiling the Clinical Incapabilities: A Benchmarking Study of GPT-4V(ision) for Ophthalmic Multimodal Image Analysis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.