Evaluation of large language model chatbot responses to psychotic prompts
Shen, E.; Hamati, F.; Donohue, M. R.; Girgis, R.; Veenstra-VanderWeele, J.; Jutla, A.
Show abstract
ImportanceThe large language model (LLM) chatbot product ChatGPT has accumulated 800 million weekly users since its 2022 launch. In 2025, several media outlets reported on individuals in whom apparent psychotic symptoms emerged or worsened in the context of using ChatGPT. As LLM chatbots are trained to align with user input, they may have difficulty responding to psychotic content. ObjectiveTo assess whether ChatGPT can reliably generate appropriate responses to prompts containing psychotic symptoms. DesignA cross-sectional study of ChatGPT responses to psychotic and control prompts, with blind clinician ratings of response appropriateness. SettingChatGPT web application accessed on 8/28-8/29/2025, testing three product versions: GPT-5 Auto (current paid default), GPT-4o (previous paid default), and "Free" (version accessible without subscription or account). Main Outcomes and MeasuresWe presented 158 unique prompts (79 control and 79 psychotic, generated based on the Structured Interview for Psychosis-Risk Syndromes) to three product versions, yielding 474 prompt-response pairs. Blinded clinicians assigned each an appropriateness rating (0 = completely appropriate, 1 = somewhat appropriate, 2 = completely inappropriate) via a standardized rubric. We hypothesized a priori that psychotic prompts would be more likely than control prompts to elicit less appropriate responses both across and within product versions. ResultsIn the primary (across-version) analysis, psychotic prompts were 25.84 times more likely to elicit less appropriate responses with "Free" ChatGPT (95% CI 12.45 to 53.66, p < 0.001). GPT-5 Auto reduced risk somewhat (OR for interaction term 0.33, 95% CI 0.16 to 0.68, p = 0.005) yet still generated less appropriate responses at a greatly elevated rate (implied OR 8.53, 95% CI 3.05 to 23.84). In the secondary (within-version) analysis, ORs were 9.08 for GPT-5 Auto (95% CI 4.24 to 21.02), 14.15 for GPT-4o (95% CI 6.12 to 37.23) and 43.37 for "Free" (95% CI 18.44 to 112.80). In an exploratory analysis, prompts reflecting grandiosity or disorganized communication were more likely to elicit inappropriate responses than those reflecting delusions. Conclusions and RelevanceNo tested version of ChatGPT reliably generated appropriate responses to psychotic content. Brief AbstractThe large language model (LLM) chatbot product ChatGPT has accumulated 800 million weekly users since its 2022 launch. In 2025, several media outlets reported on individuals in whom apparent psychotic symptoms emerged or worsened in the context of using ChatGPT. As LLM chatbots are trained to align with user input, they may have difficulty responding to psychotic content. To assess whether ChatGPT can reliably generate appropriate responses to prompts containing psychotic symptoms, we conducted a cross-sectional study of ChatGPT responses to psychotic and control prompts, with blind clinician ratings of response appropriateness. We tested three ChatGPT product versions: GPT-5 Auto (current paid default), GPT-4o (previous paid default), and "Free" (version accessible without subscription or account), presenting each with 158 unique prompts (79 control and 79 psychotic, created based on the Structured Interview for Psychosis-Risk Syndromes), yielding 474 prompt-response pairs. Blinded clinicians assigned each an appropriateness rating (0 = completely appropriate, 1 = somewhat appropriate, 2 = completely inappropriate) via a standardized rubric. We hypothesized a priori that psychotic prompts would be more likely than control prompts to elicit less appropriate responses both across and within product versions. We found that psychotic prompts were 25.84 times more likely to elicit less appropriate responses with "Free" ChatGPT (95% CI 12.45 to 53.66, p < 0.001). GPT-5 Auto reduced risk somewhat (OR for interaction term 0.33, 95% CI 0.16 to 0.68, p = 0.005) yet still generated less appropriate responses at a greatly elevated rate (implied OR 8.53, 95% CI 3.05 to 23.84). No tested version of Chat-GPT reliably generated appropriate responses to psychotic content. Key PointsO_ST_ABSQuestionC_ST_ABSCan the popular large language model product ChatGPT reliably generate appropriate responses to prompts containing psychotic content? FindingsPsychotic prompts were 26 times more likely than control prompts to elicit less appropriate responses from the current free version of ChatGPT, and 9 times more likely to elicit them from the current paid version. MeaningNo tested version of ChatGPT can reliably generate appropriate responses to psychotic content.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating the Clinical Feasibility of an Artificial Intelligence-Powered Clinical Decision Support System: A Longitudinal Feasibility Study 93%
- Evidence for feasibility of mobile health and social media-based interventions for early psychosis and clinical high risk 93%
- Development of the NeuroFlow Severity Score and Comparison With Validated Measures for Depression and Anxiety 92%
Similar papers in this journal
- The Benefits and Harms of Open Notes in Mental Health: A Delphi Survey of International Experts 93%
- Study protocol for a randomized clinical pilot trial investigating feasibility and efficacy of augmenting a virtual reality-assisted intervention targeting auditory verbal hallucinations with biofeedback: the Neuro-VR study 92%
- Shared decision-making interventions in the choice of antipsychotic prescription in people living with psychosis (SHAPE): protocol for a realist review 92%
Similar papers in this journal
- Who Does What to Whom? Graph Representations of Action-Predication in Speech Relate to Psychopathological Dimensions of Psychosis 94%
- Progressive changes in descriptive discourse in First Episode of Schizophrenia: A longitudinal computational semantics study 90%
- The efficacy of transcranial magnetic stimulation (TMS) for negative symptoms in schizophrenia: A systematic review and meta-analysis 89%
Similar papers in this journal
- Applications of Large Language Models in Psychiatry: A Systematic Review 94%
- Patients with affective disorders profit most from telemedical treatment: Evidence from a naturalistic patient cohort during the COVID-19 pandemic 92%
- Development of Goal Management Training + (GMT + ) for Methamphetamine Use Disorder Through Collaborative Design: A Process Description 92%
Similar papers in this journal
- Validation of an ICD-code-based case definition for psychotic illness across three health systems 93%
- Latent Factors of Language Disturbance and Relationships to Quantitative Speech Features 93%
- Can we detect the undetected? Comparing the prodromes of individuals with first episode psychosis detected and undetected by clinical high risk for psychosis services: an electronic health record study 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.