Extraction of Glaucoma Diagnosis, Type, and Severity from Clinical Notes using Secure Cloud-based Large Language Models
Samico, G. A.; Solages, N.; Scherer, R.; Muralidhar, R.; Gutkind, N. E.; Palazoni, V.; Medeiros, F. A.; Swaminathan, S. S.
Show abstract
Purpose: To evaluate the performance of secure cloud-based large language models (LLMs) in extracting glaucoma diagnosis, type, and severity from free-text clinical notes in the electronic health record (EHR). Design: Retrospective chart review analysis. Participants: 1,250 subjects from the Bascom Palmer Ophthalmic Repository. Methods: Clinical notes of glaucoma-related encounters between 2014 and 2024 were extracted from the Bascom Palmer Ophthalmic Repository. Two fellowship-trained glaucoma specialists annotated clinical notes for glaucoma presence, type, and severity at the eye level. The dataset was split into development (10%), validation (10%), and test (80%) sets. Development and validation sets were used for prompt engineering and refinement, and the held-out test set was used for evaluation. Five LLMs (Claude Opus 4.6, DeepSeek-V3.2, GPT-5.2, Grok 4.1, and Qwen3.6-35B-A3B) were accessed via Azure AI Foundry within HIPAA-compliant containers. Model performance was assessed using standard metrics. Clinician-entered ICD-10 codes were also compared with adjudicated labels. Main Outcome Measures: Gwet AC1, accuracy, sensitivity, specificity, and F1-score. Results: Inter-grader agreement was high for glaucoma detection (Gwet AC1= 0.930 (95% CI: 0.917-0.945), type classification (Gwet AC1= 0.917 (95% CI: 0.904-0.930), and severity staging (Gwet AC1= 0.901 (95% CI: 0.884-0.916). For glaucoma diagnosis, LLMs demonstrated high overall accuracy, with Claude achieving 97.5%, DeepSeek 96.0%, GPT 96.2%, Grok 94.4%, and Qwen 95.5%. F1 scores for glaucoma detection ranged from 95.4% to 98.9% across models. For glaucoma type classification, accuracies were 97.1%, 94.2%, 94.2%, 94.0%, and 94.4% for Claude, DeepSeek, GPT, Grok, and Qwen, respectively. F1 scores for the most prevalent type (POAG) ranged from 96.3% to 98.9%. For severity staging, accuracies were 95.0%, 94.8%, 94.5%, 94.0%, and 95.2%, respectively, with F1 scores ranging from 89.7% to 96.3% across severity categories and models. ICD-10 codes demonstrated substantially lower performance for type and severity staging, with overall accuracies of 89.2% and 58.5%, respectively. Conclusions: Secure cloud-based LLMs accurately extracted glaucoma diagnosis, type, and severity information from free-text ophthalmology notes, achieving performance approaching expert clinician adjudication while substantially outperforming ICD-based phenotyping approaches, particularly for disease severity classification. These findings demonstrate the potential of LLMs to transform unstructured clinical documentation into scalable, research-ready phenotypic data for large-scale glaucoma cohort development and EHR-based ophthalmic research.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identification of Risk Factors for Glaucoma Progression in Free-Text Clinical Notes using a Local Small Language Model 97%
- Relating Standardized Automated Perimetry Performed with Stimulus Sizes III and V in Eyes With Field Loss due to Glaucoma and NAION 95%
- Rates of Glaucoma Progression Derived from Linear Mixed Models Using Varied Random Effect Distributions 95%
Similar papers in this journal
- Artificial intelligence to facilitate clinical trial recruitment in age-related macular degeneration 94%
- Using electronic health record data to determine the safety of aqueous humor liquid biopsies for molecular analyses 94%
- Autonomous screening for Diabetic Macular Edema using deep learning processing of retinal images 94%
Similar papers in this journal
- Performance of DeepSeek-R1 in Ophthalmology: An Evaluation of Clinical Decision-Making and Cost-Effectiveness 96%
- Automated Expert-level Scleral Spur Detection and Quantitative Biometric Analysis on the ANTERION Anterior Segment OCT System 96%
- Unveiling the Clinical Incapabilities: A Benchmarking Study of GPT-4V(ision) for Ophthalmic Multimodal Image Analysis 95%
Similar papers in this journal
- Equity-Enhanced Glaucoma Progression Prediction from OCT with Knowledge Distillation 95%
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 90%
- Clinically Informed Semi-Supervised Learning Improves Disease Annotation and Equity from Electronic Health Records: A Glaucoma Case Study 90%
Similar papers in this journal
- Evaluation of OCT biomarker changes in treatment-naive neovascular AMD using a deep semantic segmentation algorithm 93%
- Dense Optic Nerve Head Deformation Estimated using CNN as a Structural Biomarker of Glaucoma Progression 93%
- An Open-Source Dataset Of Anti-Vegf Therapy In Diabetic Macular Oedema Patients Over Four Years & Their Visual Outcomes 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.