Exploring the Potential of Large Language Models for Automated Safety Plan Scoring in Outpatient Mental Health Settings
Donnelly, H. K.; Brown, G. K.; Green, K. L.; Vurgun, U.; Hwang, S.; Schriver, E.; Steinbereg, M.; Reilly, M. E.; Mehta, H.; Labouliere, C.; Oquendo, M.; Mandell, D.; Mowery, D.
Show abstract
The Safety Planning Intervention (SPI) produces a plan to help manage patients suicide risk. High-quality safety plans - that is, those with greater fidelity to the original program model - are more effective in reducing suicide risk. We developed the Safety Planning Intervention Fidelity Rater (SPIFR), an automated tool that assesses the quality of SPI using three large language models (LLMs)--GPT-4, LLaMA 3, and o3-mini. Using 266 deidentified SPI from outpatient mental health settings in New York, LLMs analyzed four key steps: warning signs, internal coping strategies, making environments safe, and reasons for living. We compared the predictive performance of the three LLMs, optimizing scoring systems, prompts, and parameters. Results showed that LLaMA 3 and o3-mini outperformed GPT-4, with different step-specific scoring systems recommended based on weighted F1-scores. These findings highlight LLMs potential to provide clinicians with timely and accurate feedback on SPI practices, enhancing this evidence-based suicide prevention strategy.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Optimising supervised machine learning algorithms predicting cigarette cravings and lapses for a smoking cessation just-in-time adaptive intervention (JITAI) 93%
- Protocol for the Houston Hospital-Based Violence Intervention Program 93%
- Suicide prevention curriculum development for health and social care students: Protocol for a scoping review 93%
Similar papers in this journal
- Artificial Intelligence (AI)-based Chatbots in Promoting Health Behavioral Changes: A Systematic Review 94%
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 93%
- Tracking private WhatsApp discourse about COVID-19: A longitudinal infodemiology study in Singapore 92%
Similar papers in this journal
Similar papers in this journal
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 92%
- Assessment of machine learning algorithms in national data to classify the risk of self-harm among young adults in hospital: a retrospective study 92%
- Completion of electronic nursing documentation of inpatient admission assessment: insights from Australian metropolitan hospitals 91%
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 91%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 91%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.