MAYOCTransformer: Masked-Attention for Yielding Comprehensive Semantic Segmentation of Retinal Optical Coherence Tomography Images using Transformer-based Neural Networks
Ye, R. Z.; Krivit, J.; Reiter, G.; Iezzi, R.
Show abstract
PurposeOptical coherence tomography (OCT) is a widely used imaging modality in ophthalmology. Accurate semantic segmentation of these images is critical for both clinical and research applications, yet existing convolutional neural network (CNN)-based methods face challenges in generalizability and robustness. This study introduces MAYOCTransformer, the first transformer-based deep learning model for comprehensive semantic segmentation of OCT images, and evaluates its performance against CNN-based models. MethodsA large dataset of 3,500 OCT images was manually segmented using an iterative deep learning-assisted workflow. The MAYOCTransformer model, based on the Mask2Former architecture, was trained and compared against CNN-based segmentation models, including U-Net, U-Net++, FPN, and DeepLabV3+. Comprehensive segmentation tasks included 10 retinal layer segmentation, choroid stroma and vessel segmentation, and the identification of 9 types of discrete pathological findings including intraretinal fluid (IRF), subretinal fluid (SRF), pigment epithelial detachment (PED), subretinal hyper-reflective material (SHRM), intraretinal hyper-reflective foci, and reticular pseudodrusen. Model performance was evaluated using the Dice similarity coefficient (DSC) on a hold-out test set with five-fold cross-validation. Additional validation was performed using external datasets, open-source segmentation models, and a randomized blinded expert evaluation. ResultsMAYOCTransformer outperformed CNN-based models in most segmentation tasks. Choroid segmentation performance was comparable between MAYOCTransformer and CNN models. External validation demonstrated the models generalizability, achieving higher DSC scores than publicly available segmentation models. A blinded expert evaluation showed that MAYOCTransformers segmentation was non-inferior to manual annotations. ConclusionMAYOCTransformer provides improved segmentation performance over CNN-based models. Its ability to generalize to external datasets suggests potential applicability in clinical and research settings.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Vis-OCT Explorer: an open-source software for visible-light optical coherence tomography data processing 96%
- A deep learning network for parallel self-denoising and segmentation in visible light optical coherence tomography of human retina 96%
- Leveraging Pretrained Vision Transformers for Automated Cancer Diagnosis in Optical Coherence Tomography Images 95%
Similar papers in this journal
- Diagnosis of central serous chorioretinopathy by deep learning analysis of en face images of choroidal vasculature 95%
- Glaucoma Detection and Staging from Visual Field Images using Machine Learning Techniques 95%
- Prediction of the ectasia screening index from raw Casia2 volume data for keratoconus identification by using convolutional neural networks 95%
Similar papers in this journal
- Interpretable Detection of Epiretinal Membrane from Optical Coherence Tomography with Deep Neural Networks 97%
- Circular functional analysis of OCT data for precise identification of structural phenotypes in the eye 95%
- Detecting Glaucoma Worsening Using Optical Coherence Tomography Derived Visual Field Estimates 94%
Similar papers in this journal
Similar papers in this journal
- Automatic wound detection and size estimation using deep learning algorithms 92%
- Explainable AI-Driven Diagnosis Model for Early Glaucoma Detection Using Grey-Wolf Optimized Extreme Learning Machine Approach 91%
- Pre-training artificial neural networks with spontaneous retinal activity improves motion prediction in natural scenes 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.