Pre-trained Vision Transformer With Masked Autoencoder for Automated Diabetic Macular Edema Detection from Optical Coherence Tomography Images
Takinami, S.; Morikawa, S.; Oshika, T.
Show abstract
PurposeTo develop and evaluate a novel self-supervised learning approach using Masked Autoencoder (MAE) pre-trained Vision Transformer (ViT) for automated detection of diabetic macular edema (DME) from optical coherence tomography (OCT) images, addressing the critical need for scalable screening solutions in diabetic eye care. Study DesignArtificial intelligence model training. MethodsWe utilized the publicly available Kermany dataset containing 109,312 OCT images, defining DME detection as a binary classification task (11,559 DME vs. 97,753 non-DME images). Five deep learning architectures were compared: MAE-pretrained ViT (MAE_ViT), standard ViT, ResNet18, VGG19_bn, and EfficientNetV2. MAE_ViT underwent two-stage training: (1) self-supervised pre-training with 75% patch masking for 1,000 epochs to learn robust visual representations, and (2) supervised fine-tuning for DME classification. Model performance was evaluated using accuracy, sensitivity, specificity, F1 score, and area under the receiver operating characteristic curve (AU-ROC) with 95% confidence intervals calculated via bootstrap resampling. ResultsMAE_ViT achieved superior performance with AU-ROC 0.999 (95% CI: 0.999-1.000), accuracy 98.5% (95% CI: 97.7-99.2%), sensitivity 99.6% (95% CI: 98.7-100%), and specificity 98.1% (95% CI: 97.2-99.1%). VGG19_bn showed the second-best performance (AU-ROC 0.997), while ResNet18 demonstrated poor specificity (28.3%) despite perfect sensitivity. The self-supervised approach of MAE_ViT outperformed standard supervised ViT (AU-ROC 0.995), demonstrating the effectiveness of learning from unlabeled data. ConclusionMAE pre-trained Vision Transformer establishes a new benchmark for automated DME detection, offering exceptional diagnostic accuracy and potential for deployment in resource-constrained settings through reduced annotation requirements.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Interpretable Detection of Epiretinal Membrane from Optical Coherence Tomography with Deep Neural Networks 97%
- ARA: accurate, reliable and active histopathological image classification framework with Bayesian deep learning 95%
- Circular functional analysis of OCT data for precise identification of structural phenotypes in the eye 94%
Similar papers in this journal
Similar papers in this journal
- Autonomous screening for Diabetic Macular Edema using deep learning processing of retinal images 95%
- A Computational Framework for Intraoperative Pupil Analysis in Cataract Surgery 94%
- Quantification of Fundus Autofluorescence Features in a Molecularly Characterized Cohort of More Than 3500 Inherited Retinal Disease Patients from the United Kingdom 92%
Similar papers in this journal
- An Inherently Interpretable AI model improves Screening Speed and Accuracy for Early Diabetic Retinopathy 97%
- Self-supervised contrastive learning improves machine learning discrimination of full thickness macular holes from epiretinal membranes in retinal OCT scans 97%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 93%
Similar papers in this journal
- Dense Optic Nerve Head Deformation Estimated using CNN as a Structural Biomarker of Glaucoma Progression 96%
- Evaluation of OCT biomarker changes in treatment-naive neovascular AMD using a deep semantic segmentation algorithm 94%
- An Open-Source Dataset Of Anti-Vegf Therapy In Diabetic Macular Oedema Patients Over Four Years & Their Visual Outcomes 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.