Prompt Engineering Limitations: Preliminary Evaluation of Large Language Models for Psychotherapy Safety
Ngo, N.; Dao, G.; Sano, A.
Show abstract
Large Language Models are increasingly used in consumer-facing mental health tools, many of which claim that prompt engineering alone can ensure safe therapeutic behavior. This study evaluates that assumption by testing 20 proprietary and open-source LLMs on high-risk psychiatric scenarios, using prompts grounded in behavioral therapy principles. Prompt engineering reduced some predictable risks, such as explicit endorsement of self-harm, but consistently failed in ambiguous or clinically nuanced situations. Models frequently validated harmful statements, colluded with hallucinations, minimized symptoms, or used stigmatizing language, including in the newest and largest models. These failures reflect structural limitations such as lack of memory, insufficient contextual reasoning, and training-related biases. Prompt engineering alone is therefore insufficient for safe AI-mediated psychotherapy; clinician-guided fine-tuning, integrated safety mechanisms, and system-level oversight will be required. This work provides early evidence motivating deeper clinician-led evaluation and safety-oriented model development.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 92%
Similar papers in this journal
- Applications of Large Language Models in Psychiatry: A Systematic Review 92%
- Development of Goal Management Training + (GMT + ) for Methamphetamine Use Disorder Through Collaborative Design: A Process Description 90%
- Understanding Psychiatric Illness Through Natural Language Processing (UNDERPIN): Rationale, Design, and Methodology 90%
Similar papers in this journal
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 94%
- The development of a World Health Organization transdiagnostic chatbot intervention for distressed adolescents and young adults 91%
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 91%
Similar papers in this journal
- Clinical code sets and the problem of redundancy in code set repositories 91%
- Optimising supervised machine learning algorithms predicting cigarette cravings and lapses for a smoking cessation just-in-time adaptive intervention (JITAI) 91%
- The Benefits and Harms of Open Notes in Mental Health: A Delphi Survey of International Experts 91%
Similar papers in this journal
- Evaluating the Clinical Feasibility of an Artificial Intelligence-Powered Clinical Decision Support System: A Longitudinal Feasibility Study 93%
- Design and Formative Evaluation of a Voice-based Virtual Coach for Problem-Solving Treatment 91%
- Development and use analysis of ‘gestioemocional.cat’, a web app for promoting emotional self-care and access to professional mental health resources during the covid-19 pandemic 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.