Automated Disease Activity Assessment in Systemic Lupus Erythematosus Using Privacy-Preserving Large Language Models
Zhang, D.; Leung, R. L.; Wong, C.-K.; Chan, S. C. W.; Li, Y.; Tang, E. H. M.; Wu, T.; Chan, T. M.; Lau, C.-S.; Wong, C. K. H.; Leung, K. S. M.; Wong, Z. S.-Y.; Wu, J. T.-K.; Yap, D. Y.-H.
Show abstract
The Systemic Lupus Erythematosus Disease Activity Index 2000 (SLEDAI-2K) is a crucial but labor-intensive tool for managing SLE. We developed a privacy-preserving, model-agnostic large language model (LLM) framework to automate SLEDAI-2K assessment from real-world electronic health records. The framework was developed on a specialist-verified ground truth of 658 clinical notes and externally validated on 56 MIMIC-IV discharge summaries. Seven open-source LLMs were evaluated using advanced prompting and ensemble strategies. The top-performing model, a two-layered GPT-OSS-120B + verifier, achieved a micro-F1 of 94.2% for descriptor classification and an 86% exact match for SLEDAI-2K scores on the internal set, with corresponding external validation performance of 87.7% and 64%, respectively. To demonstrate clinical utility, the LLMs were deployed on 2,576 serial notes from 108 SLE patients. Patients identified by the LLMs as achieving sustained low disease activity had a significantly lower incidence of stage 3 chronic kidney disease (log-rank p = 0.0053), the need for kidney replacement therapy (p = 0.044), and hospitalization (p = 0.021) over 18.3 years of follow-up. These findings demonstrate that privacy-preserving LLMs, when guided by a well-designed framework, can assist in specialist-level reasoning in autoimmune diseases, offering a scalable solution for clinical decision support and patient management.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 95%
- Conformal prediction enables disease course prediction and allows individualized diagnostic uncertainty in multiple sclerosis 94%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 93%
Similar papers in this journal
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 93%
- Consistent Performance of GPT-4o in Rare Disease Diagnosis Across Nine Languages and 4967 Cases 92%
- Machine learning guided association of adverse drug reactions with in vitro target-based pharmacology 92%
Similar papers in this journal
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 95%
- Machine learning approach to dynamic risk modeling of mortality in COVID-19: a UK Biobank study 94%
- Multiple instance learning with pathology foundation models effectively predicts kidney disease diagnosis and clinical classification 93%
Similar papers in this journal
Similar papers in this journal
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 95%
- The Interpretable Multimodal Machine Learning (IMML) framework reveals pathological signatures of distal sensorimotor polyneuropathy 92%
- A user-friendly tool for cloud-based whole slide image segmentation, with examples from renal histopathology 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.