Structured large language model extraction of clinical factors from electronic health record text supports scalable psychiatric severity prediction
Stephenson, C.; Camassa, A.; Wagner, M.; Shirazi, A. H.; Alavi, N.; Omrani, M.
Show abstract
BackgroundMental health systems face escalating demand that exceeds clinician capacity, making accurate severity-based triage a critical bottleneck. Severity assessment guides treatment intensity, resource allocation, and risk management, yet most clinically relevant information remains embedded in unstructured electronic health record (EHR) narratives, limiting its utility for scalable decision support. ObjectivesThis study evaluates whether a single large language model (LLM) can autonomously extract clinical factors from psychiatric EHR narratives, derive predictive weights from those factors, and use the resulting structured representation to predict clinician-implied severity at scale. MethodsFrom a Mayo Clinic repository of more than 2.7 million encounters, 15,000 de-identified psychiatric notes were sampled into a 5,000-patient discovery cohort and a 10,000-patient replication cohort. The same LLM (Llama 3 8B Instruct) extracted 17 background clinical factors and 3 treatment-action factors from each note. Severity reference labels were derived from the treatment-action factors using pre-specified clinical criteria. The LLM independently derived two factor-weight dictionaries from the discovery cohort: one capturing risk-oriented predictors of severe presentations and one capturing protective predictors. Five weighting conditions were then evaluated against the severity labels: the two LLM-derived dictionaries, two controls (LLM-derived variables with randomized weights; clinically irrelevant variables with arbitrary weights), and an unweighted zero-shot baseline. Performance was assessed across 928 valid iterations in the replication cohort. ResultsLLM-derived structured conditions significantly outperformed all controls and the baseline, with statistically equivalent performance between the two structured conditions. Improvements in precision and recall were balanced, indicating gains in discriminative capacity rather than threshold shifts. The variables and weights the LLM derived as predictors of severe presentations aligned closely with established clinical determinants of psychiatric severity. ConclusionA single LLM can derive clinically meaningful factor weights from unstructured EHR narratives and use them to predict psychiatric severity at scale, supporting a viable path toward interpretable, scalable triage in resource-constrained mental health systems.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting involuntary admission following inpatient psychiatric treatment using machine learning trained on electronic health record data 94%
- Relapse prevention through health technology program reduces hospitalization in schizophrenia 91%
- Antipsychotic Polypharmacy and Adverse Drug Reactions Among Adults in a London Mental Health Service, 2008-2018 91%
Similar papers in this journal
- Differential Treatment Benefit Prediction For Treatment Selection in Depression: A Deep Learning Analysis of STAR*D and CO-MED Data 94%
- Computational Mechanisms of Approach-Avoidance Conflict Predictively Differentiate Between Affective and Substance Use Disorders 93%
- Aberrant perception of environmental volatility during social learning in emerging psychosis 91%
Similar papers in this journal
Similar papers in this journal
- The path toward generalizable clinical prediction models 94%
- Receiving information on machine learning-based clinical decision support systems in psychiatric services increases staff trust in these systems: A randomized survey experiment 92%
- Psychosis Prognosis Predictor: A Continuous and Uncertainty-Aware Prediction of Treatment Outcome in First-Episode Psychosis 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.