Back

Prompt Engineering Limitations: Preliminary Evaluation of Large Language Models for Psychotherapy Safety

Ngo, N.; Dao, G.; Sano, A.

2026-07-18 psychiatry and clinical psychology
10.64898/2026.07.16.26358261 medRxiv
Show abstract

Large Language Models are increasingly used in consumer-facing mental health tools, many of which claim that prompt engineering alone can ensure safe therapeutic behavior. This study evaluates that assumption by testing 20 proprietary and open-source LLMs on high-risk psychiatric scenarios, using prompts grounded in behavioral therapy principles. Prompt engineering reduced some predictable risks, such as explicit endorsement of self-harm, but consistently failed in ambiguous or clinically nuanced situations. Models frequently validated harmful statements, colluded with hallucinations, minimized symptoms, or used stigmatizing language, including in the newest and largest models. These failures reflect structural limitations such as lack of memory, insufficient contextual reasoning, and training-related biases. Prompt engineering alone is therefore insufficient for safe AI-mediated psychotherapy; clinician-guided fine-tuning, integrated safety mechanisms, and system-level oversight will be required. This work provides early evidence motivating deeper clinician-led evaluation and safety-oriented model development.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
npj Digital Medicine
118 papers in training set
Top 0.4%
15.0%
2
PLOS ONE
5266 papers in training set
Top 22%
7.8%
3
Frontiers in Psychiatry
87 papers in training set
Top 0.2%
7.8%
4
Psychological Medicine
88 papers in training set
Top 0.2%
7.2%
5
Frontiers in Digital Health
24 papers in training set
Top 0.2%
5.4%
6
BMC Psychiatry
25 papers in training set
Top 0.2%
3.4%
7
Psychiatry Research
41 papers in training set
Top 0.4%
3.2%
8
JAMA Psychiatry
15 papers in training set
Top 0.1%
3.1%
50% of probability mass above
9
Acta Psychiatrica Scandinavica
10 papers in training set
Top 0.1%
3.1%
10
Computational Psychiatry
12 papers in training set
Top 0.1%
2.6%
11
Communications Medicine
113 papers in training set
Top 1%
2.4%
12
Nature Medicine
125 papers in training set
Top 1%
2.4%
13
The British Journal of Psychiatry
23 papers in training set
Top 0.3%
2.1%
14
Translational Psychiatry
260 papers in training set
Top 3%
1.7%
15
BMJ Mental Health
15 papers in training set
Top 0.3%
1.4%
16
JMIR Formative Research
33 papers in training set
Top 0.9%
1.4%
17
Biological Psychiatry: Cognitive Neuroscience and Neuroimaging
71 papers in training set
Top 1%
1.3%
18
Journal of Medical Internet Research
87 papers in training set
Top 2%
1.1%
19
Acta Neuropsychiatrica
14 papers in training set
Top 0.4%
1.1%
20
JMIRx Med
32 papers in training set
Top 2%
1.0%
21
BJPsych Open
29 papers in training set
Top 0.6%
1.0%
22
Schizophrenia
21 papers in training set
Top 0.3%
1.0%
23
BMC Medicine
176 papers in training set
Top 4%
1.0%
24
Nature Protocols
33 papers in training set
Top 0.5%
0.8%
25
PLOS Computational Biology
1863 papers in training set
Top 20%
0.8%
26
Frontiers in Artificial Intelligence
20 papers in training set
Top 0.8%
0.8%
27
eLife
5828 papers in training set
Top 65%
0.8%
28
Epidemiology and Psychiatric Sciences
11 papers in training set
Top 0.4%
0.8%
29
Scientific Reports
3612 papers in training set
Top 79%
0.6%
30
European Psychiatry
11 papers in training set
Top 0.4%
0.6%