Back

Structured large language model extraction of clinical factors from electronic health record text supports scalable psychiatric severity prediction

Stephenson, C.; Camassa, A.; Wagner, M.; Shirazi, A. H.; Alavi, N.; Omrani, M.

2026-05-13 psychiatry and clinical psychology
10.64898/2026.05.11.26352839 medRxiv
Show abstract

BackgroundMental health systems face escalating demand that exceeds clinician capacity, making accurate severity-based triage a critical bottleneck. Severity assessment guides treatment intensity, resource allocation, and risk management, yet most clinically relevant information remains embedded in unstructured electronic health record (EHR) narratives, limiting its utility for scalable decision support. ObjectivesThis study evaluates whether a single large language model (LLM) can autonomously extract clinical factors from psychiatric EHR narratives, derive predictive weights from those factors, and use the resulting structured representation to predict clinician-implied severity at scale. MethodsFrom a Mayo Clinic repository of more than 2.7 million encounters, 15,000 de-identified psychiatric notes were sampled into a 5,000-patient discovery cohort and a 10,000-patient replication cohort. The same LLM (Llama 3 8B Instruct) extracted 17 background clinical factors and 3 treatment-action factors from each note. Severity reference labels were derived from the treatment-action factors using pre-specified clinical criteria. The LLM independently derived two factor-weight dictionaries from the discovery cohort: one capturing risk-oriented predictors of severe presentations and one capturing protective predictors. Five weighting conditions were then evaluated against the severity labels: the two LLM-derived dictionaries, two controls (LLM-derived variables with randomized weights; clinically irrelevant variables with arbitrary weights), and an unweighted zero-shot baseline. Performance was assessed across 928 valid iterations in the replication cohort. ResultsLLM-derived structured conditions significantly outperformed all controls and the baseline, with statistically equivalent performance between the two structured conditions. Improvements in precision and recall were balanced, indicating gains in discriminative capacity rather than threshold shifts. The variables and weights the LLM derived as predictors of severe presentations aligned closely with established clinical determinants of psychiatric severity. ConclusionA single LLM can derive clinically meaningful factor weights from unstructured EHR narratives and use them to predict psychiatric severity at scale, supporting a viable path toward interpretable, scalable triage in resource-constrained mental health systems.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
npj Digital Medicine
118 papers in training set
Top 0.2%
28.3%
2
Psychological Medicine
88 papers in training set
Top 0.5%
4.3%
3
JAMA Network Open
130 papers in training set
Top 0.9%
3.5%
4
Psychiatry Research
41 papers in training set
Top 0.4%
3.5%
5
JAMIA Open
42 papers in training set
Top 0.5%
3.2%
6
Acta Neuropsychiatrica
14 papers in training set
Top 0.1%
3.2%
7
Acta Psychiatrica Scandinavica
10 papers in training set
Top 0.1%
2.6%
8
Frontiers in Psychiatry
87 papers in training set
Top 0.8%
2.6%
50% of probability mass above
9
The British Journal of Psychiatry
23 papers in training set
Top 0.2%
2.6%
10
JAMA Psychiatry
15 papers in training set
Top 0.1%
2.4%
11
Translational Psychiatry
260 papers in training set
Top 2%
2.4%
12
Frontiers in Digital Health
24 papers in training set
Top 0.6%
2.4%
13
PLOS ONE
5266 papers in training set
Top 45%
2.1%
14
Communications Medicine
113 papers in training set
Top 2%
2.1%
15
Nature Medicine
125 papers in training set
Top 1%
2.1%
16
European Psychiatry
11 papers in training set
Top 0.1%
1.7%
17
Journal of Medical Internet Research
87 papers in training set
Top 2%
1.5%
18
PLOS Medicine
110 papers in training set
Top 2%
1.5%
19
BioData Mining
22 papers in training set
Top 0.3%
1.5%
20
Biological Psychiatry: Cognitive Neuroscience and Neuroimaging
71 papers in training set
Top 0.9%
1.5%
21
Scientific Reports
3612 papers in training set
Top 60%
1.4%
22
Journal of the American Medical Informatics Association
71 papers in training set
Top 2%
1.3%
23
American Journal of Psychiatry
24 papers in training set
Top 0.4%
1.1%
24
Biological Psychiatry
137 papers in training set
Top 2%
1.1%
25
JMIR Medical Informatics
18 papers in training set
Top 0.7%
1.1%
26
BMC Psychiatry
25 papers in training set
Top 0.7%
1.1%
27
Journal of Affective Disorders
92 papers in training set
Top 1%
1.1%
28
BMC Medical Informatics and Decision Making
43 papers in training set
Top 2%
1.0%
29
Nature Neuroscience
252 papers in training set
Top 4%
1.0%
30
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 42%
0.8%