Back

Personalizing Suicide Risk Assessment: Machine Learning Extraction of Cross-Modal Interactions Between Psychosocial and Demographic Factors in Veterans

Levis, M. E.; Shiner, B.; Dimambro, M.; Rozema, L.; Ayandeh, S.; Diallo, A. B.; Zhou, Y.; Li, S.; Wu, W.; Gui, J.; Levy, J. J.

2026-06-18 psychiatry and clinical psychology
10.64898/2026.06.16.26355796 medRxiv
Show abstract

Background: Veterans face an elevated risk of suicide compared to the general population, motivating national efforts to develop predictive models that can guide proactive care. Current models used by the U.S. Department of Veterans Affairs (VA) rely primarily on structured electronic health record (EHR) data, though clinical notes contain rich contextual information that can be quantified using natural language processing (NLP) to derive psychosocial variables that may improve risk detection. Machine learning methods, particularly classification and regression trees (CART), can also uncover interactions between clinical and psychosocial variables, enabling identification of patient characteristics that modify suicide risk factors. However, integrating structured and unstructured data presents challenges because NLP features often greatly outnumber traditional clinical variables, potentially biasing interaction discovery. In prior work, we addressed this imbalance by introducing a weighted CART framework that balances structured variables with NLP-derived psychosocial features from semantic lexicons (SEANCE). While effective, semantic approaches summarize language into predefined constructs and may overlook important lexical variation present in clinical narratives. Methods: In this study, we extend that framework by replacing semantic features with a high-dimensional bag-of-words (BoW) representation of clinical notes and by evaluating models across cohorts defined by structured suicide risk stratification (low, medium, high) and varying temporal lookback windows. Using a cohort of 27,241 veterans, we analyzed clinical documentation collected up to 30, 90, or 270 days prior to death (or a matched index date for controls), enabling temporally flexible risk modeling. XGBoost models were trained to balance structured and unstructured features and identify cross-modal interactions between textual and clinical variables. Results: When incorporated into generalized linear models, these interactions improved predictive performance, particularly among low- and medium-risk patients, and substantially reduced the performance gap between interpretable and more complex models. Notably, the BoW representation outperformed our prior semantic index-based approach. Discussion and Conclusions: Together, these findings demonstrate the utility of interpretable NLP methods for uncovering clinically meaningful interactions between psychosocial and demographic factors in suicide risk and establish a strong benchmark for future deep learning approaches aimed at capturing richer contextual and temporal information from clinical narratives.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
npj Digital Medicine
118 papers in training set
Top 0.3%
22.0%
2
Frontiers in Artificial Intelligence
20 papers in training set
Top 0.1%
10.7%
3
BioData Mining
22 papers in training set
Top 0.1%
9.7%
4
JAMIA Open
42 papers in training set
Top 0.3%
5.5%
5
PLOS ONE
5266 papers in training set
Top 35%
3.5%
50% of probability mass above
6
Journal of Medical Internet Research
87 papers in training set
Top 0.7%
3.4%
7
BMC Medical Informatics and Decision Making
43 papers in training set
Top 0.6%
3.2%
8
Journal of Affective Disorders
92 papers in training set
Top 0.7%
3.2%
9
Psychiatry Research
41 papers in training set
Top 0.5%
2.8%
10
Scientific Reports
3612 papers in training set
Top 42%
2.5%
11
Acta Neuropsychiatrica
14 papers in training set
Top 0.1%
2.4%
12
Frontiers in Digital Health
24 papers in training set
Top 0.8%
1.7%
13
JMIR Medical Informatics
18 papers in training set
Top 0.5%
1.5%
14
PLOS Computational Biology
1863 papers in training set
Top 17%
1.1%
15
European Psychiatry
11 papers in training set
Top 0.2%
1.1%
16
Biological Psychiatry: Cognitive Neuroscience and Neuroimaging
71 papers in training set
Top 1%
1.1%
17
Psychological Medicine
88 papers in training set
Top 1%
1.1%
18
Journal of the American Medical Informatics Association
71 papers in training set
Top 2%
1.1%
19
Translational Psychiatry
260 papers in training set
Top 3%
1.1%
20
Nature Medicine
125 papers in training set
Top 3%
0.9%
21
JMIRx Med
32 papers in training set
Top 2%
0.8%
22
Life
29 papers in training set
Top 0.9%
0.8%
23
American Journal of Epidemiology
67 papers in training set
Top 1%
0.8%
24
Acta Psychiatrica Scandinavica
10 papers in training set
Top 0.2%
0.8%
25
Pharmacoepidemiology and Drug Safety
18 papers in training set
Top 0.5%
0.8%
26
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 41%
0.8%
27
Frontiers in Psychiatry
87 papers in training set
Top 2%
0.8%
28
Communications Medicine
113 papers in training set
Top 6%
0.6%
29
JMIR Public Health and Surveillance
45 papers in training set
Top 2%
0.6%