An independent supervisory safety agent improves reaction of large language models to suicidal ideation
Trivedi, S.; Simons, N. W.; Tyagi, A.; Ramaswamy, A.; Nadkarni, G. N.; Charney, A. W.
Show abstract
BackgroundLarge language models (LLMs) are increasingly used in mental health contexts, yet their detection of suicidal ideation is inconsistent, raising patient safety concerns. MethodsWe conducted a cross-sectional evaluation using 224 paired suicide-related clinical vignettes presented in a single-turn format under two conditions (with and without structured clinical information). Native LLM safeguard responses were compared with an independent supervisory safety architecture with asynchronous monitoring. The primary outcome was detection of suicide risk requiring intervention. ResultsThe supervisory system detected suicide risk in 205 of 224 evaluations (91.5%) versus 41 of 224 (18.3%) for native LLM safeguards. Among 168 discordant evaluations, 166 favored the supervisory system and 2 favored the LLM (matched odds ratio {approx}83.0). Both systems detected risk in 39 evaluations, and neither in 17. Detection was highest in scenarios with explicit suicidal ideation and lower in more ambiguous presentations. ConclusionsNative LLM safeguards frequently failed to detect suicide risk in this structured evaluation. An independent monitoring approach substantially improved detection, supporting the role of external safety systems in high-risk mental health applications of LLMs.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Applications of Large Language Models in Psychiatry: A Systematic Review 92%
- Development of Goal Management Training + (GMT + ) for Methamphetamine Use Disorder Through Collaborative Design: A Process Description 91%
- Co-development of a best practice checklist for mental health data science: A Delphi study 90%
Similar papers in this journal
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 92%
- The development of a World Health Organization transdiagnostic chatbot intervention for distressed adolescents and young adults 91%
- Precision Digital Intervention for Depression Based on Social Rhythm Principles Adds Significantly to Outpatient Treatment 90%
Similar papers in this journal
- Optimising supervised machine learning algorithms predicting cigarette cravings and lapses for a smoking cessation just-in-time adaptive intervention (JITAI) 91%
- A One-Arm Pilot Trial of a Telehealth CBT-Based Group Intervention Targeting Transdiagnostic Risk for Emotional Distress 90%
- The Benefits and Harms of Open Notes in Mental Health: A Delphi Survey of International Experts 90%
Similar papers in this journal
- Integrating Expert Knowledge into Large Language Models Improves Performance for Psychiatric Reasoning and Diagnosis 92%
- Piloting Forensic Tele-Mental Health Evaluations of Asylum Seekers 91%
- Symptom Monitoring based on Digital Data Collection During Inpatient Treatment of Schizophrenia Spectrum Disorders – a Feasibility Study 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.