Assessment of the Modified Rankin Scale in Electronic Health Records with a Fine-tuned Large Language Model
Silva, L.; Milani, M.; Bindra, S.; ikramuddin, S. S.; Tessmer, M.; Frederickson, K.; Datta, A.; Ergen, H.; Stangebye, A.; Cooper, D.; Kumar, K.; Yeung, J.; Lakshminarayan, K.; Streib, C.
Show abstract
IntroductionThe modified Rankin scale (mRS) is an important metric in stroke research, often used as a primary outcome in clinical trials and observational studies. The mRS can be assessed retrospectively from electronic health records (EHR), though this process is labor-intensive and prone to inter-rater variability. Large language models (LLMs) have demonstrated potential in automating clinical text classification. We hypothesize that a fine-tuned LLM can analyze EHR text and classify mRS scores for clinical and research applications. MethodsWe performed a retrospective cohort study of patients admitted to a specialist stroke neurology service at a large academic hospital system between August 2020 and June 2023. Each patients medical record was reviewed at two time points: (1) hospital discharge and (2) approximately 90 days post-discharge. Two independent researchers assigned an mRS score at each time point. Two separate models were trained on EHR passages with corresponding mRS scores as labeled outcomes: (1) a multiclass model to classify all seven mRS scores and (2) a binary model to classify functional independence (mRS 0-2) versus non-independence (mRS 3-6). Four-fold cross-validation was conducted, using accuracy and Cohens kappa as model performance metrics. ResultsA total of 2,290 EHR passages with corresponding mRS scores were included in model training. The multiclass model--considering all seven scores of the mRS--attained an accuracy of 77% and a weighted Cohens Kappa of 0.92. Class-specific accuracy was highest for mRS 4 (90%) and lowest for mRS 2 (28%). The binary model--considering only functional independence vs non-independence --attained an accuracy of 92% and Cohens Kappa of 0.84. ConclusionOur findings demonstrate that LLMs can be successfully trained to determine mRS scores through EHR text analysis. With further advancements, fully automated LLMs could scale across large clinical datasets, enabling data-driven public health strategies and optimized resource allocation.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- EchoGraph System for Automated Quality Assessment of Echocardiography Reports 93%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 92%
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 92%
Similar papers in this journal
- Use of unstructured text in prognostic clinical prediction models: a systematic review 94%
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 93%
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 93%
Similar papers in this journal
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 93%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 92%
- A proposed de-identification framework for a cohort of children presenting at a health facility in Uganda 91%
Similar papers in this journal
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 94%
- ARDSFlag: An NLP/Machine Learning Algorithm to Visualize and Detect High-Probability ARDS Admissions Independent of Provider Recognition and Billing Codes 92%
- An Informatics Consult approach for generating clinical evidence for treatment decisions 92%
Similar papers in this journal
- Preliminary outcomes of combined treadmill and overground high-intensity interval training in ambulatory chronic stroke 92%
- Automated Identification of Thrombectomy Amenable Vessel Occlusion on Computed Tomography Angiography using Deep Learning 91%
- Workflow Intervals andOutcomesof Endovascular Treatment for Acute Large-Vessel Occlusion During On- Versus Off-Hours in China The ANGEL-ACT Registry 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.