Evaluating Chatbots in Psychiatry: Rasch-Based Insights into Clinical Knowledge and Reasoning
Chang, Y.; Huang, S.-S.; Hsu, W.-Y.; Liu, Y.-C.
Show abstract
Chatbots are increasingly being recognized as valuable tools for clinical support in psychiatry. This study systematically evaluates their strengths and limitations in psychiatric clinical knowledge and reasoning. A total of 27 chatbots, including ChatGPT-o1-preview, were assessed using 160 multiple-choice questions derived from the 2023 and 2024 Taiwan Psychiatry Licensing Examinations. The Rasch model was employed to analyze chatbot performance, supplemented by dimensionality analysis and qualitative assessments of reasoning processes. Among the models, ChatGPT-o1-preview achieved the highest performance, with a JMLE ability score of 2.23, significantly exceeding the passing threshold (p < 0.001). It excelled in diagnostic and treatment reasoning and demonstrated a strong grasp of psychopharmacology concepts. However, limitations were identified in its factual recall, handling of niche topics, and occasional reasoning biases. Building on these findings, we have highlighted key aspects of a potential clinical workflow to guide the practical integration of chatbots into psychiatric practice. While ChatGPT-o1-preview holds significant potential as a clinical decision-support tool, its limitations underscore the necessity of human oversight. Continuous evaluation and domain-specific training are crucial to maximize its utility and ensure safe clinical implementation.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Applications of Large Language Models in Psychiatry: A Systematic Review 96%
- Patients with affective disorders profit most from telemedical treatment: Evidence from a naturalistic patient cohort during the COVID-19 pandemic 93%
- Development of Goal Management Training + (GMT + ) for Methamphetamine Use Disorder Through Collaborative Design: A Process Description 93%
Similar papers in this journal
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 95%
- Validation of Visual and Auditory Digital Markers of Suicidality in Acutely Suicidal Psychiatric In-Patients 92%
- Artificial Intelligence (AI)-based Chatbots in Promoting Health Behavioral Changes: A Systematic Review 91%
Similar papers in this journal
- The Benefits and Harms of Open Notes in Mental Health: A Delphi Survey of International Experts 93%
- Psychotherapies and Psychological Support for Individuals Facing Psychological Distress during the COVID-19 Pandemic: A Scoping Review 93%
- Study protocol for a randomized clinical pilot trial investigating feasibility and efficacy of augmenting a virtual reality-assisted intervention targeting auditory verbal hallucinations with biofeedback: the Neuro-VR study 92%
Similar papers in this journal
- Evaluating the Clinical Feasibility of an Artificial Intelligence-Powered Clinical Decision Support System: A Longitudinal Feasibility Study 94%
- Development of the NeuroFlow Severity Score and Comparison With Validated Measures for Depression and Anxiety 93%
- Development and use analysis of ‘gestioemocional.cat’, a web app for promoting emotional self-care and access to professional mental health resources during the covid-19 pandemic 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.