Translational Vision Science & Technology
● Association for Research in Vision and Ophthalmology (ARVO)
All preprints, ranked by how well they match Translational Vision Science & Technology's content profile, based on 39 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Chia, Z. K.; Kong, A. W.; Turner, M. L.; Saifee, M.; Damato, B. E.; Backus, B. T.; Blaha, J. J.; Schuman, J. S.; Deiner, M. S.; Ou, Y.
Show abstract
ObjectiveTo assess the feasibility of remotely training glaucoma patients to take a ten-session clustered virtual reality (VR) visual field test (VVP-10) at home, analyze results for test-retest variability, and assess correspondence with conventional perimetry. DesignCross-sectional study. Subjects21 subjects with glaucoma were enrolled and included in the feasibility assessment of remote training. 36 eyes were used for test-retest analysis and determination of concordance with Humphrey Visual Field (HVF) testing. MethodsSubjects were provided with a mobile VR headset containing the VVP-10 test software and trained remotely via video conferencing. Subjects were instructed to complete ten sessions over a 14-day period. Main Outcome MeasuresFeasibility was determined by the number of subjects who were able to independently complete VVP-10 over the 14-day period after one remote training session. Intraclass correlation coefficient (ICC) of average fraction seen across ten sessions and standard error (SE) for the mean were primary outcome measures for assessing test-retest variability. Correlation with HVF mean sensitivity (MS) across eyes, was a secondary outcome measure. Results20 subjects (95%) successfully completed the VVP-10 test series after one training session. ICC of VVP-10 was 0.95 (95% CI [0.92, 0.97]). Mean SE in units of fraction seen was 0.012. The Spearman correlations of VVP-10 average fraction seen versus HVF MS were 0.88 (95% CI [0.66, 0.99]) for moderate to advanced glaucoma eyes, and decreased to 0.68 (95% CI [0.29, 0.94]) when all eyes were included. ConclusionsRemote training of patients at home is feasible and subsequent remote clustered visual field testing using VVP-10 by patients on their own without any further interactions with caregivers or study staff was possible. At-home VVP-10 results demonstrated low test-retest variability. Future studies must be conducted to determine if VVP-10, taken at home as convenient for the patient, may be a viable supplement to provide equivalent or complementary results to that of standard in-clinic assessment of visual function in glaucoma.
Vrijling, A. C. L.; de Boer, M. J.; Renken, R. J.; Marsman, J.-B. C.; Heutink, J.; Cornelissen, F. W.; Jansonius, N. M.
Show abstract
PurposeStandard automated perimetry (SAP) is the gold standard for functional assessment in glaucoma. SAP can be too demanding for some groups of patients. Continuous visual stimulus tracking (SONDA: Standardized Oculomotor and Neurological Disorders Assessment) simplifies the perimetric task to following a moving stimulus on a screen. In this study we evaluated the screening performance of SONDA-based eye movement perimetry (SONDA-EMP) in glaucoma. To explore generalizability, we evaluated an experimental setup (SONDA-Eyelink) and a clinic-ready version (SONDA-Neon). MethodsSONDA-Eyelink and SONDA-Neon measurements were performed in 100 cases with glaucoma (36, 36, and 28 with early, moderate, and severe glaucoma, respectively) and 100 age-similar controls. Participants monocularly tracked a moving stimulus (Goldmann size III) at 40% contrast (both setups) and 160% (SONDA-Eyelink). Eye movements were continuously recorded. Outcome was the agreement between gaze and stimulus position. We used previously collected glaucoma case-control data to build a continuous glaucoma screening score. This score was used for an ROC-analysis applied to the current, independently collected dataset. We predefined good screening performance as: at 95% specificity, a sensitivity of at least 50%, 90%, and 100% for early, moderate, and severe glaucoma, respectively. ResultsAt 95% specificity, the sensitivity of SONDA-Eyelink was 58, 94, and 100% at 40% contrast and 56, 97, and 100% at 160% contrast for early, moderate, and severe glaucoma, respectively. Sensitivity was 53, 94, and 100% for SONDA-Neon. ConclusionsSONDA-EMP is a novel, fast, and intuitive method to screen for visual function loss in glaucoma.
Diaz, J. C.; Skerswetat, J.; Browne, A. W.
Show abstract
PurposeTo determine the repeatability of a clinical pupillometer in healthy participants with a 5 and 30-minute test-retest break for both Swinging-Flashlight and Low-High Luminance approach. Methods20 healthy participants (mean age: 27 years {+/-}5; 9 females) placed their heads into the devices headrest and fixated a central spot. A Swinging-Flashlight approach evoked a pupillary light reflex, stimulating alternatingly 8 times each eye by a brief diffuse white light flash followed by a continuous measurement of constriction and dilation of the direct and consensual pupils. For the Low-High Luminance setting, continuous measurements of pupil diameters for a 5 second low luminance display followed by a 5 second high luminance white light. Pupillary light reflex parameters for each test and each eye were calculated by the device. Repeatability was investigated after a 5-minute break time and during a control experiment for 21 participants after a 30-minute break for each approach in counterbalanced order using Bland-Altman analysis. ResultsOverall, many parameters for both Swinging-Flashlight and Low-High Luminance approach showed retest biases for all pupil light reflex parameters after a 5-minute test-retest break. These biases were almost completely reduced after a 30-minute break between test and retest for both approaches. The test-retest variabilities as expressed using the Coefficient of Repeatability was reduced after 30-minutes for the majority of the Low-High Luminance results but not for results of the Swinging-Flashlight method. ConclusionsClinical pupillometry includes the Swinging Flashlight Test (SFL) and Low-High Luminance assays. In SFL, many pupillary dynamic metrics showed test-retest bias at a 5-minute interval, whereas relative afferent pupillary defect (RAPD) measurements remained unbiased at both 5 and 30-minute intervals. Similarly, Low-High Luminance testing showed minimal bias after 30 minutes but exhibited bias at the 5-minute interval. Thus, when evaluating multiple pupillary parameters in scientific or clinical settings, RAPD is less susceptible to bias, while other metrics require longer intervals between testing and retesting.
Koyama, M.; Inoda, S.; Ueno, Y.; Ito, Y.; Oshika, T.; Tanito, M.
Show abstract
PurposeTo train and evaluate segmentation-free 3D convolutional neural network (3DCNN) models for estimating visual field (VF) from optical coherence tomography (OCT) images and to independently assess the longitudinal variability and progression detection capabilities of Humphrey Field Analyzer (HFA) measurements and OCT-based estimated VF (OCT-VF) in a diverse clinical population. DesignRetrospective multicenter study. Participants13,366 patients (24,313 eyes) underwent HFA tests (24-2, trimmed 30-2, or 10-2 test patterns) and macular OCT imaging at five ophthalmic institutions. The dataset included 129,007 paired OCT-VF data points representing various ocular conditions. MethodsWe trained segmentation-free 3DCNN models using comprehensive OCT datasets without disease-specific exclusions, employing 10-fold cross-validation to estimate VF thresholds and mean deviation (MD). Unlike previous studies, we independently assessed both OCT-VF and HFA measurements by creating separate longitudinal datasets with standardized measurement counts and observation periods for comparative analysis, enabling direct evaluation of clinical reliability. We analyzed absolute residual variability from regression lines using jackknife resampling, applied Bonferroni correction for multiple comparisons, and used Spearmans correlation for progression analysis. Main Outcome MeasuresOCT-VF and HFA VF agreement, residual variability, progression detection rates, and progression rate correlations. ResultsOCT-VF and HFA VF showed correlations (Pearsons r: 24-2 thresholds 0.863, MD 0.924; 10-2 thresholds 0.881, MD 0.939; all p < 0.001). OCT-VF demonstrated significantly lower residual variability than HFA for all parameters (OCT-VF vs. HFA: 0.58 vs. 1.12 dB for 24-2 MD; 0.70 vs. 1.12 dB for 10-2 MD; all p < 0.001). This advantage persisted across all test points (mean variability reduction: 60.4% for 24-2; 55.1% for 10-2), age groups, and most severity levels. OCT-VF identified more progression events (24-2 MD: 113% more, 10-2 MD: 48.6% more). MD slopes showed correlations between OCT-VF and HFA (Pearsons r: 24-2 MD 0.831, 10-2 MD 0.863; all p < 0.001). ConclusionsThe segmentation-free 3DCNN models objectively estimated VF from OCT images with significantly lower longitudinal variability than performance-dependent HFA measurements across diverse ocular conditions. The lower variability of OCT-VF enhances statistical power for progression detection, suggesting its clinical potential as a complementary tool derived from routine OCT data, decreasing measurement noise, and enabling more timely therapeutic interventions.
MUQIT, M. M.; LE MER, Y.; OLMOS DE KOO, L. C.; HOLZ, F. G.; SAHEL, J. A.; Palanker, D.
Show abstract
ObjectiveTo assess the efficacy and safety of the PRIMA subretinal neurostimulation system 48-months post-implantation for improving visual acuity (VA) in patients with geographic atrophy (GA) due to age-related macular degeneration (AMD) at 48-months post-implantation. DesignFirst-in-human clinical trial of the PRIMA subretinal prosthesis in patients with atrophic AMD, measuring best-corrected ETDRS VA (Clinicaltrials.gov NCT03333954). SubjectsFive patients with GA, no foveal light perception and VA of logMAR 1.3 to 1.7 in their worse-seeing "study" eye. MethodsIn patients implanted with a subretinal photovoltaic neurostimulation array containing 378 pixels of 100 m in size, the VA was measured with and without the PRIMA system using ETDRS charts at 1 meter. The systems external components: augmented reality glasses and pocket computer, provide image processing capabilities, including zoom. Main Outcome MeasuresVA using ETDRS charts with and without the system. Light sensitivity in the central visual field, as measured by Octopus perimetry. Anatomical outcomes demonstrated by fundus photography and optical coherence tomography up to 48-months post- implantation. ResultsAll five subjects met the primary endpoint of light perception elicited by the implant in the scotoma area. In one patient the implant was incorrectly inserted into the choroid. One subject died 18-months post-implantation due to study-unrelated reason. ETDRS VA results for the remaining three subjects are reported herein. Without zoom, VA closely matched the pixel size of the implant: 1.17 {+/-} 0.13 pixels, corresponding to mean logMAR 1.39, or Snellen 20/500, ranging from 20/438 to 20/565. Using zoom at 48 months, subjects improved their VA by 32 ETDRS letters versus baseline (SE 5.1) 95% CI[13.4,49.9], p<0.0001. Natural peripheral visual function in the treated eye did not decline after surgery compared to the fellow eye (p=0.08) during the 48 months follow-up period. ConclusionsSubretinal implantation of PRIMA in subjects with GA suffering from profound vision loss due to AMD is feasible and well tolerated, with no reduction of natural peripheral vision up to 48-months. Using prosthetic central vision through photovoltaic neurostimulation, patients reliably recognized letters and sequences of letters,and with zoom it provided a clinically meaningful improvement in VA of up to eight ETDRS lines.
Zhang, A.; Pratap, J. S.; Young, J. R.; Lui, J.; Attari, K.; Srivastava, A. A.; Wu, E. R.; Moreno, A. D.; Miall, A.; Iyer, J. M.; Mantena, S.; Chandra, J.; Joshi, V. P.; Jacobs, D. S.
Show abstract
PurposeFluorescein staining (FS) is a standard method of assessing corneal epithelium (CE) integrity. However, the equipment and personnel required for FS may be unavailable in low-resource environments. We developed and validated a low-cost, noninvasive, and quantitative CE evaluation pipeline using a custom smartphone attachment and convolutional neural networks (CNNs). MethodsA 3D-printed smartphone attachment and placido disk illumination module was attached to a OnePlus 7 Pro smartphone. 26 smartphone-acquired images were obtained from 15 subjects, comprising a dataset including healthy eyes and corneal epitheliopathies of Oxford grade I-V. A classifier CNN was trained on 8 subjects (23,173 image patches) to identify areas of suspected epithelial disruption, and validated on 7 subjects (10,883 image patches). The fraction of disrupted corneal surface area (FDSA) was computed for each subject from the model output. Results were compared with FS slit lamp photos which were independently graded by two clinicians using the Oxford scheme. ResultsFDSA showed promise as a non-invasive marker of CE integrity, with mean FDSA in the Oxford >II cohort being higher than the Oxford [≤]II cohort (p = 0.04 and p = 0.09 using Oxford scores from each clinician, respectively). Additionally, areas of CE disruption identified by our smartphone-based technique showed qualitative concordance with those revealed by FS. ConclusionsOur technique for smartphone-based CE imaging and automated analysis is a promising low-cost, noninvasive method to quantitatively evaluate the CE. Translational RelevanceThis tool can be used to evaluate ocular surface disease in low-resource regions.
Yang, G.; Gaffney, M.; Cooper, R. F.
Show abstract
More than 50 inherited retinal diseases are known, with rod photoreceptors serving as early indicators in nearly half of them. Given that current clinical imaging modalities cannot resolve the loss of individual rods, rod optoretinograms hold unique promise for transforming the early detection and precise monitoring of retinal disease. They provide a powerful means of detecting early functional changes in diseases marked by rod degeneration, such as Retinitis Pigmentosa, conditions like Age-related Macular Degeneration and certain forms of night blindness, where rod mosaics remain structurally intact, but physiologically impaired. The "optoretinogram" is a relatively new assay that operates through detection of optical changes in cells in response to stimuli. This tool has excellent potential for providing insights into the earliest functional changes of individual photoreceptors, with the potential to assist in the early detection, monitoring, and treatment of retinal diseases. In this work, we obtained intensity-based optoretinograms (iORGs) from rod photoreceptors using an adaptive optics scanning laser ophthalmoscope. We explore the necessity of both individual rod identification for extracting these waveforms and discuss rod iORG RMS morphology in the context of previously reported cone iORGs. We found that human rod iORG RMS waveforms have slower implicit times and lower amplitudes than cone iORG RMS waveforms. Additionally, we determined that we obtain very similar iORG RMS metrics using either only rod locations or all pixels where rods reside. The ability to obtain rod optoretinograms without counting individual rods greatly simplifies the functional evaluation of rods and makes the approach more practical and scalable for larger populations and diseased retina.
Parikh, K. S.; Reddy, K.; Shuff, J.; Kong, X.; Li, X.; Hariharakumar, S.; Santhanaraj, V. A.; Kakoty, R.; Kamam, B.; Ravilla, P. K.; Vivekanand, N.; Hoopes, M.; Detels, K.; Mohseni, N.; Verma, R.; Shumeyko, D.; Yadalla, D.; Venkatesh, R.; Shekhawat, N. S.
Show abstract
BackgroundCataract and anterior segment diseases are leading causes of blindness in low-resource settings. Eye camp screenings remain the primary mode of community outreach but are constrained by cost, logistics, and dependence on highly trained specialists. We designed and validated a low-cost, user-friendly smartphone-based anterior segment imaging and teleophthalmology platform to enable community health workers (CHWs) to perform diagnostic-quality eye screening. MethodsWe designed a portable imaging device (ScoutTM) paired with an accessible Android smartphone and mobile application (InSightful) for telemedicine in low-bandwidth settings. CHWs underwent 3 hours of training on using the imaging software and hardware, then screened patients across 19 rural eye camps in South India with anterior segment images and clinical data uploaded to a cloud-based database for remote ophthalmologist (RO) review. Diagnoses and referral decisions made by ROs were compared with those of in-person eye camp ophthalmologists (ECOs). CHWs, ROs, and patients were surveyed on the platforms feasibility and acceptability. FindingsN=1093 patients underwent eye camp screening by ECOs and CHW-led smartphone screening with RO review. CHWs completed screenings in <2.5 minutes/eye and obtained diagnostic-quality images for >90% of eyes. ROs and ECOs showed 96.1% concordance for referral decisions (95% CI 94.7-97.1) and substantial agreement in diagnosis of any cataract ({kappa} 0.77, percent agreement [PA] 89%), mature cataract ({kappa} 0.67, PA 96%), immature cataract ({kappa} 0.69, PA 85%), clear crystalline lens ({kappa} 0.65, PA 89%), pseudophakia ({kappa} 0.92, PA 97%), and moderate agreement for pterygium ({kappa} 0.47, PA 94%). Concordance increased with image quality. CHWs, ROs, and patients reported high usability, acceptability, and net promoter scores. InterpretationScoutTM anterior segment screening by minimally trained CHWs achieves diagnostic and referral accuracy comparable to in-person ophthalmologist examinations, supporting potential to decentralize cataract screening and expand access to eye care in low-resource settings. FundingNational Eye Institute R21EY034343, National Eye Institute K23EY032988, National Eye Institute P30EY01765 (Biostatistics Core), Microsoft Innovation Acceleration Award, Johns Hopkins Center for Global Health, Stephen F Raab and Mariellen Brickley-Raab Rising Professorship in Ophthalmology, Boone Pickens Rising Professorship in Ophthalmology
Chuter, B.; Huynh, J.; Walker, E.; Hallaj, S.; Jalili, J.; Liebmann, J.; Fazio, M. A.; Girkin, C. A.; Weinreb, R. N.; Christopher, M.; Zangwill, L. M.
Show abstract
PurposeTo fine tune and evaluate the performance of the retinal foundation model (RETFound) on a diverse longitudinal clinical research dataset in glaucoma detection from optical coherence tomography (OCT) RNFL scans. Subanalyses of the model performance were evaluated across different subgroups, various dataset sample sizes and training cycles (epochs). DesignEvaluation of a diagnostic technology Subjects, Participants, and Controls15,216 Spectralis OCT RNFL circle scans of 747 individuals of diverse race (56.9% White, 37.8% Black/African American, and 5.3% Other/Not reported, glaucoma severity (30.8% mild, 18.4% moderate-to-severe, and 50.9% no glaucoma), and age (44.8% <60 years, 55.2% >60 years) from the Diagnostic Innovations in Glaucoma Study (DIGS) and the African Descent and Glaucoma Evaluation Study (ADAGES). All OCT scans were labeled as "Non-glaucomatous" or "Glaucomatous." MethodsRETFound was employed to perform binary glaucoma classification. The diagnostic accuracy of RETFound was iteratively tested across different combinations of dataset sample sizes (50 to 2000 OCT RNFL circle scans), epochs (5 to 50), and study subpopulations stratified by severity of glaucoma, age, and race). Main Outcome MeasuresArea under receiver operating characteristic curve (AUC) for classifying RNFL scans as "Non-glaucomatous" or "Glaucomatous." ResultsPerformance metrics improved with larger training datasets and more training cycles, rising from an AUC of 0.61 (50 training images and 5 epochs) to AUC 0.91 (2,000 training images and 50 epochs). Gains in performance were marginal as training size increased beyond 500 scans. Performance was similar across race for all training size and cycle number combinations: African American (AUC=0.90) vs other (AUC=0.93). RNFL scans from older patients (>60 years) led to worse performance (AUC=0.85) compared to younger patients (<60 years, AUC=0.95). Performance was significantly higher for RNFL scans from patients with moderate-to-severe glaucoma vs mild glaucoma (AUC=0.99 vs 0.88, respectively). ConclusionsGood RETFound performance was observed with a relatively small sample size of images used for fine tuning and across differences in race and age. RETFounds ability to adapt across a range of OCT training conditions and populations suggests it is a promising tool to automate glaucoma detection in a variety of use cases. PrecisThe study found high accuracy for glaucoma detection from OCT optic nerve head RNFL scans in a diverse study population by adapting an existing foundation model (RETFound). Performance improved with larger datasets and more training cycles, achieving an AUC of 0.91 with RNFL scans alone. Results suggest RETFound is promising for automated OCT RNFL-based glaucoma detection across demographics and training conditions.
Banijamali, S. M. A.; Versek, C. W.; Lashkari, K.; Bex, P.; Sridhar, S.
Show abstract
PurposeDelayed Dark-Adapted vision Recovery (DAR) is a known biomarker for Age-related Macular Degeneration (AMD); however, its measurement is often cumbersome for both patients and examiners. In this study, we developed NeuroVEP, a portable, wireless, and user-friendly system designed to objectively assess Dark-Adapted Visual Evoked Potentials (DAVEP). MethodsNeuroVEP consists of a headset with a smartphone that delivers controlled photo-bleach and monocular pattern reversal stimuli while utilizing custom electroencephalography (EEG) electrodes and electronics to measure DAVEP. The system allows for separate analysis of the near peripheral and macular visual field of each eye, completing the test in a comfortable, single-session format (<25 minutes) without requiring subjective patient feedback. The NeuroVEP test protocol included: (i) Mesopic luminance pattern reversal VEP for macular and peripheral regions (5 mins), (ii) Full-field photopic pattern reversal VEP (2.5 mins), (iii) Scotopic luminance DAVEP recovery post photo-bleach (up to 15 mins), measured simultaneously from both eyes. The data were analyzed for 66 participants, divided into four cohorts: (A) Age-matched healthy controls with no ophthalmic pathologies (n=10), (B) Early-stage AMD (AREDS1) (n=19), (C) Intermediate-stage AMD (AREDS3) (n=18), (D) Advanced-stage AMD (AREDS4/5) (n=19). Advanced signal processing and machine learning methodologies were applied to filter and process the VEP responses from the DAR segment of the experiment. 13 discriminating features were extracted from the processed signals and classified for each participant using a Bayesian statistical framework and Gaussian Mixture Model (GMM). ResultsThe algorithm demonstrated: 86% accuracy in early-stage AMD detection (Healthy vs. Early AMD) (Sensitivity: 97%, Specificity: 65%, AUC-ROC: 0.81 and AUC-PR: 0.92) and 93% accuracy in overall AMD detection (Healthy vs. All AMD stages) (Sensitivity: 98%, Specificity: 65%, AUC-ROC: 0.82 and AUC-PR: 0.97). ConclusionsWe successfully developed a portable, objective user-friendly VEP system and an advanced Bayesian-GMM statistical analysis framework capable of identifying DAR deficits in AMD patients. This novel technology shows high potential for early AMD detection and could serve as a non-invasive, objective diagnostic tool for AMD screening in clinical and remote settings.
Nguyen, V.; Iyengar, S.; Rasheed, H.; Apolo, G.; Li, Z.; Kumar, A.; Nguyen, H.; Bohner, A.; Dhodapkar, R.; Do, J.; Duong, A.; Gluckstein, J.; Hong, K.; James, A.; Lee, J.; Nguyen, K.; Wong, B.; Ambite, J.-L.; Kesselman, C.; Daskivich, L.; Pazzani, M.; Xu, B.
Show abstract
PurposeTo develop and test a deep learning (DL) algorithm for detecting referable glaucoma in the Los Angeles County (LAC) Department of Health Services (DHS) teleretinal screening program. MethodsFundus photographs and patient-level labels of referable glaucoma (defined as cup-to-disc ratio [CDR] [≥] 0.6) provided by 21 trained optometrist graders were obtained from the LAC DHS teleretinal screening program. A DL algorithm based on the VGG-19 architecture was trained using patient-level labels generalized to images from both eyes. Area under the receiver operating curve (AUC), sensitivity, and specificity were calculated to assess algorithm performance using an independent test set that was also graded by 13 clinicians with one to 15 years of experience. Algorithm performance was tested using reference labels provided by either LAC DHS optometrists or an expert panel of 3 glaucoma specialists. Results12,098 images from 5,616 patients (2,086 referable glaucoma, 3,530 non-glaucoma) were used to train the DL algorithm. In this dataset, mean age was 56.8 {+/-} 10.5 years with 54.8% females and 68.2% Latinos, 8.9% Blacks, 2.7% Caucasians, and 6.0% Asians. 1,000 images from 500 patients (250 referable glaucoma, 250 non-glaucoma) with similar demographics (p [≥] 0.57) were used to test the DL algorithm. Algorithm performance matched or exceeded that of all independent clinician graders in detecting patient-level referable glaucoma based on LAC DHS optometrist (AUC = 0.92) or expert panel (AUC = 0.93) reference labels. Clinician grader sensitivity (range: 0.33-0.99) and specificity (range: 0.68-0.98) ranged widely and did not correlate with years of experience (p [≥] 0.49). Algorithm performance (AUC = 0.93) also matched or exceeded the sensitivity (range: 0.78-1.00) and specificity (range: 0.32-0.87) of 6 LAC DHS optometrists in the subsets of the test dataset they graded based on expert panel reference labels. ConclusionsA DL algorithm for detecting referable glaucoma developed using patient-level data provided by trained LAC DHS optometrists approximates or exceeds performance by ophthalmologists and optometrists, who exhibit variable sensitivity and specificity unrelated to experience level. Implementation of this algorithm in screening workflows could help reallocate eye care resources and provide more reproducible and timely glaucoma care.
Bolo, K.; Nguyen, T. H.; Iyengar, S.; Li, Z.; Nguyen, V.; Wong, B.; Do, J.; Ambite, J.-L.; Kesselman, C.; Daskivich, L.; Xu, B.
Show abstract
PurposeTo compare the performance of a foundation model and a supervised learning-based model for detecting referable glaucoma from fundus photographs. DesignEvaluation of diagnostic technology. Participants6,116 participants from the Los Angeles County Department of Health Services Teleretinal Screening Program. MethodsFundus photographs were labeled for referable glaucoma (cup-to-disc ratio [≥] 0.6) by certified optometrists. Four deep learning models were trained on cropped and uncropped images (Training N = 8,996; Validation N = 3,002) using two architectures: a vision transformer with self-supervised pretraining on fundus photographs (RETFound) and a convolutional neural network (VGG-19). Models were evaluated on a held-out test set (N = 1,000) labeled by glaucoma specialists and an external test set (N = 300) from University of Southern California clinics. Performance was assessed while varying training set size and stratifying by demographic factors. xRAI was used for saliency mapping. Main Outcome MeasuresArea under the receiver operating characteristic curve (AUC-ROC) and threshold-specific metrics. ResultsThe cropped image VGG-19 model achieved the highest AUC-ROC (0.924 [0.907-0.940]), which was comparable (p = 0.07) to the cropped image RETFound model (0.911 [0.892-0.930]), which achieved the highest Youden-optimal performance (sensitivity 82.6%, specificity 88.2%) and F1 score (0.801). Cropped image models outperformed their uncropped counterparts within each architecture (p < 0.001 for AUC-ROC comparisons). RETFound models had a performance advantage when trained on smaller datasets (N < 2000 images), and the uncropped image RETFound model performed best on external data (p < 0.001 for AUC-ROC comparisons). The cropped image RETFound model performed consistently across ethnic groups (p = 0.20), while the others did not (p < 0.04); performance did not vary by age or gender. Saliency maps for both architectures consistently included the optic nerve. ConclusionWhile both RETFound and VGG-19 models performed well for classification of referable glaucoma, foundation models may be preferable when training data is limited and when domain shift is expected. Training models using images cropped to the region of the optic nerve improves performance regardless of architecture but may reduce model generalizability.
Chaurasia, A. K.; Wang, C.; Toohey, P. W.; Chen, C. Y.; MacGregor, S.; Bennett, M. T.; Verma, N.; Craig, J. E.; McCartney, P. J.; Sarossy, M. G.; Hewitt, A. W.
Show abstract
BackgroundThe visual field (VF) test results of many eyes with glaucoma progress despite treatment. This suggests that some eyes are either untreated or that the management of intraocular pressure (IOP) does not influence the outcome. In this work, we explore whether future VF parameters can be predicted from a baseline optical coherence retinal nerve fibre layer (OCT-RNFL) scan using a deep learning model. MethodsThe model was developed using 1792 eyes from 1610 patients, and externally validated on 151 eyes from a second centre using the same Zeiss Cirrus machine and 281 eyes from a third centre using scans obtained from a different (Heidelberg Spectralis) machine. The Vision Transformers (ViT)-based regression model was trained on baseline OCT-RNFL scans to predict three key VF indices (follow-up interval: 4.74 {+/-} 2.59 years). Model performance was evaluated using Mean Absolute Error (MAE) and Root Mean Square Error (RMSE), with 95% confidence intervals (CI). ResultsThe model achieved an overall MAE of 2.07 (95% CI: 1.91-2.22) and RMSE of 2.87 (95% CI: 2.60-3.14) on the internal validation set. On external validation, the model showed comparable performance with an MAE of 2.07 (95% CI: 1.8-2.35) for the external validation (Zeiss OCT) cohort and 2.11 (95% CI: 1.93-2.31) for the external validation (Heidelberg OCT) cohort. Saliency maps revealed that the inner and outer RNFL layers were key structures in driving the models predictions. ConclusionsOur ViT-based regression model effectively predicts key VF indices objectively from a single OCT-RNFL scan, with strong performance across two OCT devices, offering a novel tool for predicting glaucoma progression.
Kannan, V. P.; Shuff, J.; Acharya, A.; Venkatesh, R.; Shekhawat, N. S.; Parikh, K. S.
Show abstract
Globally, cataract remains the leading cause of blindness, affecting over 100 million people, with a disproportionate burden in low- and middle-income countries where access to ophthalmologists is limited. Although cataract surgery can restore vision almost immediately, timely diagnosis and referral remains a major barrier to care. We developed and prospectively evaluated lightweight, multimodal machine learning models capable of classifying lens status on a smartphone, enabling accessible screening in low-resource environments. We trained and evaluated both early and late fusion model architectures to classify lens status as clear, immature cataract, mature cataract, or pseudophakia using 6,794 anterior segment eye images captured using Scout smartphone-based diffuse illumination system paired with clinical data (age, visual acuity, and pinhole acuity) from 2,956 patients from three eye hospitals in India. The early fusion model, which jointly integrates image and clinical features via a learnable gating mechanism, achieved superior performance (AUROC=0.98) compared to late fusion. Model interpretation using feature importance and Grad-CAM revealed that early fusion effectively balanced visual and clinical parameters, mirroring ophthalmologist diagnostic reasoning. Prospective on-device evaluation in 210 patients at the Aravind Eye Hospital demonstrated equivalent performance, achieving an AUROC of 0.96, confirming robustness and real-time feasibility on mobile hardware. These results demonstrate the first prospectively validated, on-device, multimodal cataract machine learning model, demonstrating the feasibility of instant, offline cataract classification and referral in low-resource environments. This advance has potential to broaden cataract screening, allowing minimally trained workers to screen and refer patients, and enabling earlier diagnosis, referral, and treatment in underserved populations.
Ipek-Ugay, S.; Zeyadi, G.
Show abstract
BackgroundAchieving precise postoperative refractive outcomes remains a significant challenge in cataract surgery. While advanced intraocular lens (IOL) power calculation formulas exist, they are constrained by their singular algorithmic structures. This study investigated whether a stacking ensemble machine learning approach could overcome these limitations. MethodsA dataset of 1,710 eyes from patients who underwent cataract surgery with monofocal IOL implantation (Vivinex or SA60AT) was utilized. Following rigorous preprocessing and feature engineering, a stacking ensemble architecture was developed comprising three diverse base learners (Multi-Layer Perceptron, Support Vector Regressor with RBF kernel, and SplineTransformer with Linear Regression) and a Ridge Regressor meta-learner. The model was trained on 80% of the data using 5-fold cross-validation and evaluated on an independent 20% test set (n=341). Performance was compared against six standard IOL formulas. ResultsThe stacking ensemble model demonstrated excellent predictive accuracy, achieving a Mean Absolute Error (MAE) of 0.272 D on the independent test set (n=341). The model achieved lower MAE compared to all six standard IOL formulas, including Kane (MAE 0.295 D) and Barrett Universal II (MAE 0.318 D). Clinically, 85.1% of eyes achieved predictions within {+/-}0.50 D, compared to 82.5% for Kane formula and 81.8% for Barrett Universal II. ConclusionThe stacking ensemble machine learning model significantly enhances postoperative refraction prediction accuracy compared to established IOL calculation formulas. By leveraging algorithmic diversity and data-driven learning, this approach represents a promising advancement toward reducing refractive surprises and improving patient satisfaction in cataract surgery. External validation on independent datasets is required to confirm generalizability.
Dvey-Aharon, Z.; Lalman, C.; Ianchulev, T.; Livne, M.; Margalit, D.; Aviv, R.; Mendoza, K. A. V.; Schuman, J. S.
Show abstract
PurposeGlaucoma, a leading cause of irreversible vision loss, often remains undiagnosed due to its asymptomatic progression and the limitations of existing screening methods. This study aimed to validate an artificial intelligence machine learning algorithm for the camera-agnostic detection of glaucomatous optic neuropathy using macula-centered fundus images. MethodsData were collected from EyePACS, a teleretinal screening system, comprising 25,000 macula-centered fundus images from 12,500 patients at U.S. primary care centers. A secondary dataset from the Philadelphia Telemedicine Glaucoma Follow-up Study was used for independent validation. A convolutional neural network was developed to detect glaucomatous optic neuropathy. Expert-graded fundus images served as the ground truth. Images underwent quality filtering to ensure the visibility of the optic nerve. Bilateral images were analyzed to produce patient-level diagnoses. Validation involved a secondary dataset of fundus images. ResultsThe sensitivity and specificity of the algorithm in detecting glaucomatous optic neuropathy is calculated in comparison to expert grading. From the EyePACS dataset, 21,792 images (10,986 subjects) met quality standards. The algorithm demonstrated a sensitivity of 90.6% and specificity of 90.5%. Validation on the secondary dataset (200 fundus images from 100 subjects) resulted in a sensitivity of 96.4% and specificity of 85.3%. ConclusionsThe algorithm achieved high sensitivity and specificity in detecting glaucomatous optic neuropathy using macula-centered fundus images, demonstrating its potential for integration into diverse clinical settings. Its camera-agnostic design and robust performance offer a scalable solution for improving glaucoma screening pathways, making them more accessible and efficient.
Baek, J. S.; Lokhande, A.; Neuenschwander, D.; Shi, M.; Wang, M.
Show abstract
Purpose To investigate the relative efficacy of nine distinct visual field (VF) denoising artificial intelligence (AI) methods and a pathology-aware AI strategy to discourage over-correction of glaucomatous defects. Design Retrospective study. Participants 87,940 paired visual field (VF) and optical coherence tomography (OCT) samples from a tertiary academic center. Methods Denoising models were trained on a separate VF-only dataset and evaluated on an independent structure-function dataset of paired VF-OCT samples. We implemented and evaluated nine distinct VF denoising strategies representing three broad categories: baseline measurements, self-supervised and image restoration models (including Noise2Noise, Noise2Void, and NAFNet), and latent variable compression-based models (autoencoders and variational autoencoders). All models were designed to reconstruct VF sensitivity maps. We then predicted retinal nerve fiber layer thickness (RNFLT) maps from the denoised VFs using a fixed, independently trained VF-to-RNFLT prediction model. Main Outcome Measures Predicted VF and RNFLT maps and resultant evaluation metrics. Results The raw VF baseline achieved a global R2 of 0.5468 and MAE of 16.83 um. Restoration-based models maintained or slightly improved concordance, with the pathology-aware NAFNet achieving the highest global R2 of 0.5485 and a comparable MAE of 16.82 um. In contrast, compression-based models degraded concordance, with CNN-VAE showing a significant reduction (R2 approximately 0.50). In severe glaucoma, concordance decreased across all methods; however, compression architectures exhibited disproportionately greater degradation compared with restoration-based approaches. Conclusions We present a comparative benchmark of AI-based VF denoising strategies paired with structure-function evaluation. While restoration-based models can reduce variability without loss of biological signal, latent compression risks attenuating clinically meaningful defects. Visually smoother fields are not necessarily more biologically accurate.
Moradi, M.; Cao-Xue, J.; Eslami, M.; Wang, M.; Elze, T.; Zebardast, N.
Show abstract
Forecasting glaucoma progression remains a major challenge in preventing irreversible vision loss. We developed and validated a multimodal, longitudinal deep learning framework to predict future progression using a large retrospective cohort of 10,864 patients from Mass Eye and Ear. The model integrates sequential structural (OCT RNFL scans), functional (visual-field maps), and clinical data from a two-year observation window to forecast progression over the subsequent two-to four-year horizon. Four backbone architectures (ConvNeXt-V2, ViT, MobileNet-V2, EfficientNet-B0) were coupled with a bidirectional LSTM to capture temporal dynamics. The ConvNeXt-V2-based model achieved 0.97 AUC and 0.94-0.96 accuracy, outperforming other backbones with robust performance across sex and race subgroups and only modest attenuation in those > 70 years. Saliency maps localized to clinically relevant arcuate bundles, supporting biological plausibility. By effectively fusing multimodal data over time, this framework enables accurate, interpretable, and equitable long-horizon risk stratification, advancing personalized glaucoma management.
Karaer, I.; Yoon, H.-J.; Ma, R.; Savant, R.; Rodwell, V.; Shenoy, R.; Tu, Z.; Arshad, Q.; Mukaetova-Ladinska, E. B.; Thomas, M. G.
Show abstract
Background/ObjectivesVirtual Reality (VR) eye trackers offer portable, objective tools for neuro-ophthalmic testing. This study evaluated the feasibility, reproducibility and reliability of a VR eye tracker (BulbiCAM) compared to wearable eye-tracking glasses (PupilLabs Neon glasses), highlighting its potential clinical utility and feasibility. Subjects/MethodsA prospective study involving 39 healthy participants (mean age{+/-}SD = 30.0 {+/-} 9.5 years) assessed inter-visit reproducibility of BulbiCAM tests across two visits. Pupillary light reflex tests were conducted with both BulbiCAM and PupilLabs Neon, enabling paired assessments. Reproducibility was analysed using intra-class correlation coefficients (ICC), reliability via Bland-Altman analysis, and participant experience through a survey evaluating test comfort and usability. ResultsParticipants feedback (n=27) highlighted high acceptability for BulbiCAM: 89% found the test comfortable, 92.6% felt the testing duration was appropriate, and 81.5% reported no eye strain or fatigue. Inter-visit reproducibility of Bulbicam tests showed high reproducibility for pursuit and pupil tests (ICC= 0.88-0.76), while saccadic tasks showed lower reproducibility (best ICC at 0.62). Paired assessments between devices showed close agreement for key pupillometer metrics: baseline diameter (bias: -0.48 {+/-} 0.47 mm), peak constriction diameter (bias: -0.56 {+/-} 0.36 mm), constriction velocity (bias: 0.22 {+/-} 0.58 mm/s), and duration of constriction (bias: -0.052 {+/-} 0.15 s). ConclusionsThis study highlights the clinical feasibility of BulbiCAM, with high patient acceptability and reproducibility for pursuit and pupil tests. Paired assessments confirmed its accuracy for key pupillometric parameters, validating its reliability for clinical and research use.
Lipsky, T.; Ehrenzeller, C.; Ansari, G.; Pfau, K.; Harmening, W.; Wu, Z.; Pfau, M.
Show abstract
Purpose: To quantify whether fundus tracking in microperimetry improves psychometric parameter estimation (in vivo demonstration of improved stimulus-delivery precision), and to derive a psychometrically grounded criterion intensity for suprathreshold (defect-mapping) microperimetry. Methods: Twenty-five healthy volunteers underwent MAIA2-microperimetry at five loci: three outside and two inside the blind spot. Frequency-of-seeing (FoS) functions were measured in four blocks (2 tracking on; 2 tracking off). FoS-data were fit using cumulative-Gaussian psychometric functions estimating sensitivity parameters. Mixed-effect models assessed tracking effects, and posterior simulations defined the optimal criterion intensity for separating 'seeing' from 'non-seeing' loci. Results: Tracking had little effect on threshold estimates at loci outside the blind spot, but lowered threshold estimates within the blind spot (posterior median difference PMD [95% CrI] of -1.46 dB [-2.30, -0.62] at locus 4, and -1.02 dB [-1.94, -0.08] at locus 5). Tracking was associated with steeper psychometric slope parameters at loci 1-3 (PMD of -0.14 dB [-0.29, 0.01], -0.27 dB [-0.43, -0.12], and -0.22 dB [-0.40, -0.04]). Without tracking, false-positive responses were more frequent when fixation shifts displaced stimuli toward the 'seeing' retina. Simulation-based analysis identified 13 dB as nominally optimal criterion for suprathreshold microperimetry (Youden index: 0.76 [0.74, 0.79], comparable to 10 dB (0.74 [0.72, 0.76]). Conclusions: Even in healthy volunteers with stable fixation, fundus tracking measurably reduced sensitivity estimates at 'non-seeing' loci and sharpened FoS curves in the 'seeing' retina. A criterion intensity of 10 to 13 dB is a defensible choice for separating 'seeing' and 'non-seeing' retina in suprathreshold (defect-mapping) perimetry paradigms.