OphthaBERT: Automated Glaucoma Diagnosis from Clinical Notes
Shah, R.; Moradi, M.; Eslami, S.; Fujita, A.; Aziz, K.; Bineshfar, N.; Elze, T.; Eslami, M.; Kazeminasab, S.; Liebman, D.; Rasouli, S.; Vu, D.; Wang, M.; Yohannan, J.; Zebardast, N.
Show abstract
Glaucoma is a leading cause of irreversible blindness worldwide, with early intervention often being crucial. Research into the underpinnings of glaucoma often relies on electronic health records (EHRs) to identify patients with glaucoma and their subtypes. However, current methods for identifying glaucoma patients from EHRs are often inaccurate or infeasible at scale, relying on International Classification of Diseases (ICD) codes or manual chart reviews. To address this limitation, we introduce (1) OphthaBERT, a powerful general clinical ophthalmology language model trained on over 2 million diverse clinical notes, and (2) a fine-tuned variant of OphthaBERT that automatically extracts binary and subtype glaucoma diagnoses from clinical notes. The base OphthaBERT model is a robust encoder, outperforming state-of-the-art clinical encoders in masked token prediction on out-of-distribution ophthalmology clinical notes and binary glaucoma classification with limited data. We report significant binary classification performance improvements in low-data regimes (p < 0.001, Bonferroni corrected). OphthaBERT is also able to achieve superior classification performance for both binary and subtype diagnosis, outperforming even fine-tuned large decoder-only language models at a fraction of the computational cost. We demonstrate a 0.23-point increase in macro-F1 for subtype diagnosis over ICD codes and strong binary classification performance when externally validated at Wilmer Eye Institute. OphthaBERT provides an interpretable, equitable framework for general ophthalmology language modeling and automated glaucoma diagnosis.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Artificial intelligence to facilitate clinical trial recruitment in age-related macular degeneration 94%
- Quantification of Fundus Autofluorescence Features in a Molecularly Characterized Cohort of More Than 3500 Inherited Retinal Disease Patients from the United Kingdom 92%
- Evaluating the Performance of ChatGPT in Ophthalmology: An Analysis of its Successes and Shortcomings 91%
Similar papers in this journal
- Equity-Enhanced Glaucoma Progression Prediction from OCT with Knowledge Distillation 95%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 91%
- Clinically Informed Semi-Supervised Learning Improves Disease Annotation and Equity from Electronic Health Records: A Glaucoma Case Study 91%
Similar papers in this journal
- Identification of Risk Factors for Glaucoma Progression in Free-Text Clinical Notes using a Local Small Language Model 93%
- Advancing Question-Answering in Ophthalmology with Retrieval Augmented Generations (RAG): Benchmarking Open-source and Proprietary Large Language Models 93%
- AutoMorph: Automated Retinal Vascular Morphology Quantification via a Deep Learning Pipeline 91%
Similar papers in this journal
- An Inherently Interpretable AI model improves Screening Speed and Accuracy for Early Diabetic Retinopathy 94%
- Self-supervised contrastive learning improves machine learning discrimination of full thickness macular holes from epiretinal membranes in retinal OCT scans 93%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 90%
Similar papers in this journal
- Performance of DeepSeek-R1 in Ophthalmology: An Evaluation of Clinical Decision-Making and Cost-Effectiveness 92%
- Unveiling the Clinical Incapabilities: A Benchmarking Study of GPT-4V(ision) for Ophthalmic Multimodal Image Analysis 90%
- Automated Expert-level Scleral Spur Detection and Quantitative Biometric Analysis on the ANTERION Anterior Segment OCT System 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.