Back

Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study

Kalinich, M.; Luccarelli, J.; Santa Maria, J.; Flathers, M.; Nguyen, A.; Song, S. H.; Makhoul, K.; Rivera Criado, M. J.; Ginapp, C. M.; Hill, B.; Shumate, J. N.; Notsu, H.; Smith, C.; Moss, F.; Torous, J.

2026-07-14 psychiatry and clinical psychology
10.64898/2026.07.10.26357132 medRxiv
Show abstract

Background General-purpose large language models increasingly encounter emotional and therapy-like conversation, yet are not developed or evaluated as clinical systems. Existing safety evaluations rely largely on brief exchanges, although harms often unfold over extended interactions. Whether models maintain safety-relevant performance as conversations accumulate context remains unknown. Methods In this preregistered study, 400 clinician-validated statements, with or without suicidal ideation, were inserted at 0-200 speaker turns in 5 psychotherapy and 3 synthetic transcripts. Forty-nine LLMs and 8 clinicians performed the same binary classification task. Mixed-effects models estimated the effects of conversational depth, model scale, and model version on F1. Twelve top models were tested to 1,500 turns across conversational trajectories, with or without instruction restatement. Results F1 declined with depth across model families (p<0.001). Larger, newer models performed better but still degraded. Clinicians showed no decline (mean F1 0.86 at both 0 and 200 turns), but eight of nine proprietary models exceeded their performance at 200 turns. Conversational content, not length alone, explained F1 changes; the largest decrease was under adversarial context (p<0.001). Restating instructions increased F1 on therapy to near baseline (median {Delta}F1 +0.12; p<0.001; 89% median recovery) versus MSJ ({Delta}F1 +0.08; p=0.04; 38% recovery). Conclusions LLM detection of suicidal ideation degraded with conversational depth and trajectory, whereas clinician performance remained stable despite the strongest models exceeding most clinicians in absolute performance. Mental health AI safety evaluations should test sustained performance across realistic and adversarial trajectories rather than relying on short-prompt benchmarks.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
npj Digital Medicine
118 papers in training set
Top 0.2%
26.7%
2
PLOS ONE
5266 papers in training set
Top 24%
6.8%
3
Frontiers in Digital Health
24 papers in training set
Top 0.1%
6.3%
4
Frontiers in Psychiatry
87 papers in training set
Top 0.3%
5.5%
5
Psychiatry Research
41 papers in training set
Top 0.4%
3.5%
6
JAMA Network Open
130 papers in training set
Top 1%
3.2%
50% of probability mass above
7
Psychological Medicine
88 papers in training set
Top 0.7%
3.2%
8
Acta Psychiatrica Scandinavica
10 papers in training set
Top 0.1%
1.7%
9
JAMA Psychiatry
15 papers in training set
Top 0.2%
1.7%
10
Scientific Reports
3612 papers in training set
Top 53%
1.7%
11
The British Journal of Psychiatry
23 papers in training set
Top 0.3%
1.7%
12
Nature Medicine
125 papers in training set
Top 2%
1.7%
13
Computational Psychiatry
12 papers in training set
Top 0.1%
1.7%
14
Translational Psychiatry
260 papers in training set
Top 3%
1.4%
15
Biological Psychiatry: Cognitive Neuroscience and Neuroimaging
71 papers in training set
Top 1.0%
1.4%
16
Journal of Medical Internet Research
87 papers in training set
Top 2%
1.3%
17
Acta Neuropsychiatrica
14 papers in training set
Top 0.4%
1.1%
18
BMJ Mental Health
15 papers in training set
Top 0.3%
1.1%
19
BMC Medicine
176 papers in training set
Top 4%
1.1%
20
Communications Medicine
113 papers in training set
Top 3%
1.1%
21
BMC Psychiatry
25 papers in training set
Top 0.6%
1.1%
22
JMIR Formative Research
33 papers in training set
Top 1%
1.1%
23
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 37%
1.1%
24
European Psychiatry
11 papers in training set
Top 0.2%
1.0%
25
PLOS Medicine
110 papers in training set
Top 3%
1.0%
26
JMIRx Med
32 papers in training set
Top 2%
1.0%
27
BioData Mining
22 papers in training set
Top 0.7%
0.9%
28
Frontiers in Artificial Intelligence
20 papers in training set
Top 0.7%
0.8%
29
Nature Protocols
33 papers in training set
Top 0.5%
0.8%
30
PLOS Computational Biology
1863 papers in training set
Top 21%
0.6%