Comparative Evaluation of Pretrained Large Language Models for Suicide Risk Prediction from Clinical Notes in U.S. Veterans
Levy, J.; Levis, M.; Dimambro, M.; Rozema, L.; Ayandeh, S.; Diallo, A.; Zhou, Y.; Li, S.; Wu, W.; Shiner, B.; Gui, J.
Show abstract
Background: Suicide remains a significant and potentially preventable cause of death among United States veterans. Predictive models based on structured electronic health record (EHR) data, including the U.S. Department of Veterans Affairs' Recovery Engagement and Coordination for Health-Veterans Enhanced Treatment (REACH-VET) program, aim to identify individuals at elevated risk for enhanced monitoring and follow-up. Increasing evidence suggests that unstructured clinical narratives contain additional psychosocial information that may enhance risk prediction when analyzed using natural language processing (NLP). However, optimal approaches for representing clinical text remain uncertain. Recent advances in large language models (LLMs) enable contextual text representations that capture complex semantic relationships beyond traditional lexical methods. Methods: We compared the predictive performance of pretrained LLMs with classical bag-of-words (BoW) representations for suicide risk prediction using clinical notes from 27,241 veterans receiving care in the Veterans Health Administration. Patients were stratified by REACH-VET risk tier (low, moderate, high), and models were evaluated across prediction windows defined by note look-back periods (<30, <90, and <270 days). Results: LLM-based representations outperformed BoW approaches in seven of nine risk tier-time window combinations, achieving a maximum AUROC of 0.644 when solely considering text. Incorporating structured clinical variables further improved performance (AUROC=0.748). Model interpretation identified suicide-related language, especially in notes documented within 30 days of the outcome among patients classified as high risk. Conclusions: Pretrained LLMs can extract clinically meaningful information from narrative documentation, providing a foundation for future work adapting to additional clinical contexts and nuanced temporal associations to improve suicide risk prediction.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 95%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 93%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 93%
Similar papers in this journal
- Improving ascertainment of suicidal ideation and suicide attempt with natural language processing 96%
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 94%
- Using explainable machine learning to identify patients at risk of reattendance at discharge from emergency departments 94%
Similar papers in this journal
- Psychiatric disorders and self-harm across 26 adult cancers: cumulative burden, temporal variation, excess years of life lost and unnatural causes of deaths 91%
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 90%
- Zero-shot drug repurposing with geometric deep learning and clinician centered design 90%
Similar papers in this journal
- Combining AI and human support in mental health: a digital intervention with comparable effectiveness to human-delivered care 91%
- Uncovering social states in healthy and clinical populations using digital phenotyping and Hidden Markov Models 91%
- Validation of Visual and Auditory Digital Markers of Suicidality in Acutely Suicidal Psychiatric In-Patients 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.