Development, System Design, Safety, and Performance Metrics of a Conversational Agent for Reducing Depressive and Anxious Symptoms Based on a Large Language Model: The MHAI Study
Villarreal-Zegarra, D.; Paredes-Gonzales, Y.; Damaso-Roman, A.; Quinones-Inga, J.; Centeno-Terrazas, G.; Lozada, Y. P. A.-M.
Show abstract
BackgroundConversational agents based on large language models (LLMs) have shown moderate efficacy in reducing depressive and anxiety symptoms. However, most existing evaluations lack methodological transparency, rely on closed-source models, and show limited standardization in performance and safety assessment. ObjectiveWe have two study objectives: (1) to develop an LLM-based conversational agent through system design analysis and initial functionality testing, and (2) to evaluate its safety and performance through standardized assessment in controlled simulated interactions focused on depression and anxiety of two LLMs (GPT-4o and Llama 3.1-8B). MethodsWe conducted a cross-sectional study in two phases. First, we developed a mental health platform integrating a conversational agent with functionalities including personalized context, pretrained therapeutic modules, self-assessment tools, and an emergency alert system. Second, we evaluated the agents responses in simulated interactions based on predefined user personas for each LLM. Four expert raters assessed 816 interaction pairs using a 5-criterion Likert scale evaluating tone, clarity, domain accuracy (correctness), robustness, completeness, boundaries, target language, and safety. In addition, we use performance metrics based on numerical criteria such as cost, response length, and number of tokens. Multiple linear regression models were used to compare LLM performance and assess metric interrelations. ResultsFirst, we developed a web-based mental health platform using a user-centered design, structured into frontend, backend, and database layers. The system integrates therapeutic chat (GPT-4o and Llama 3.1-8B), psychological assessments (PHQ-9, GAD-7), CBT-based tasks, and an emergency alert system. The platform supports secure user authentication, data encryption, multilingual access, and session tracking. Second, GPT-4o outperformed Llama 3.1-8B in both performance metrics based on numerical criteria and Likert scale criteria, generating longer and more lexically diverse responses, using more tokens, and scoring higher in clarity, robustness, completeness, boundaries, and target language. However, it incurred higher costs, with no significant differences in tone, accuracy, or safety. ConclusionOur study presents a conversational agent with multiple functionalities and shows that GPT-4o outperforms Llama 3.1-8B in performance, although at a higher cost. This platform could be used in future clinical trials or real-world implementation studies.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Design and Formative Evaluation of a Voice-based Virtual Coach for Problem-Solving Treatment 95%
- Development and use analysis of ‘gestioemocional.cat’, a web app for promoting emotional self-care and access to professional mental health resources during the covid-19 pandemic 95%
- Evaluating the Clinical Feasibility of an Artificial Intelligence-Powered Clinical Decision Support System: A Longitudinal Feasibility Study 94%
Similar papers in this journal
- Listening to mental health crisis needs at scale: using Natural Language Processing to understand and evaluate a mental health crisis text messaging service 94%
- The development of a World Health Organization transdiagnostic chatbot intervention for distressed adolescents and young adults 93%
- Precision Digital Intervention for Depression Based on Social Rhythm Principles Adds Significantly to Outpatient Treatment 91%
Similar papers in this journal
- Artificial Intelligence (AI)-based Chatbots in Promoting Health Behavioral Changes: A Systematic Review 95%
- Tracking private WhatsApp discourse about COVID-19: A longitudinal infodemiology study in Singapore 95%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 94%
Similar papers in this journal
- Evaluation of easy-to-implement anti-stress interventions in a series of N-of-1 trials: Study protocol of the Anti-Stress Intervention Among Physicians Study (ASIP) 93%
- Applications of Large Language Models in Psychiatry: A Systematic Review 92%
- Development of medical device software for the screening and assessment of depression severity using data collected from a wristband-type wearable device: SWIFT study protocol 92%
Similar papers in this journal
- Digital Health Tools for the Passive Monitoring of Depression: A Systematic Review of Methods 93%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 92%
- Evaluation of two easy-to-implement digital breathing interventions in the context of daily stress levels in a series of N-of-1 trials: results from the Anti-Stress Intervention Among Physicians (ASIP) Study 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.