Evaluating the Readability and Reliability of Large Language Model Generated Information About Anxiety and Depression
Porwal, G.; Laddha, K.; Jeenger, J.
Show abstract
Artificial intelligence (AI) large language models (LLMs) hold great potential to transform psychiatry and mental health care by delivering relevant and tailored mental health information. This study aimed to evaluate the quality of mental health information generated by LLMs by determining their level of accessibility, reliability, and interpreting any bias present. Generative Pre-trained Transformer-4 (GPT-4) (San Francisco, California: Open AI), Gemini 1.5 Flash (Mountain View, California: Google LLC), and Large Language Model Meta AI 3.2 (Llama 3.2) (Menlo Park, California: Meta Inc.) were prompted with 20 questions commonly asked about anxiety and depression. The responses were evaluated using Flesch-Kincaid readability tests to quantify their ease of understanding through grade level and reading ease score measures. The text was subsequently analyzed using a modern DISCERN score to assess the reliability of the health-related information presented. Finally, the LLM responses were evaluated for stigmatizing language using communication and language guidelines for mental health. A significant difference in grade levels and reading ease scores was observed between GPT-4 and Llama (p < 0.01) and Gemini and Llama(p < 0.01), with both GPT-4 and Gemini having higher readability scores. No significant difference was observed in the grade level and reading ease scores between GPT-4 and Gemini (p > 0.05). All three models reported moderate reliability scores but no significant differences were observed (p > 0.05). GPT-4, Gemini, and Llama included stigmatizing phrases in 10%, 15%, and 20% of their responses respectively. These phrases were attributed to descriptions of mental health conditions and substance use; however, no significant differences were observed in the proportions of stigmatizing phrases present across the three models (p > 0.05). Furthermore, comparisons between anxiety and depression responses for each model revealed no significant differences in readability, reliability, or bias. All models demonstrated the ability to generate mental health information that generally satisfied criteria for accessibility, reliability, and minimal bias, with GPT-4 and Gemini reporting higher readability than Llama. However, the lack of additional resources cited in responses and the occasional presence of certain stigmatizing phrases, indicate that additional model training and fine-tuning, specifically for mental health applications, may be necessary before these tools can be deployed on a large scale for mental health care.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Defining Destigmatizing Design Guidelines for Use in Sexual Health-Related Digital Technologies: A Delphi Study 94%
- Developing contents for a digital drug adherence tool with reminder cues and personalized feedback: a formative mixed-methods study among children and adolescents living with HIV in Tanzania 93%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 93%
Similar papers in this journal
- Evaluating the Clinical Feasibility of an Artificial Intelligence-Powered Clinical Decision Support System: A Longitudinal Feasibility Study 94%
- Development and use analysis of ‘gestioemocional.cat’, a web app for promoting emotional self-care and access to professional mental health resources during the covid-19 pandemic 94%
- Exploring Patient and Staff Experiences of Video Consultations During COVID-19 in an English Outpatient Care Setting: Secondary Data Analysis of Routinely Collected Feedback Data 93%
Similar papers in this journal
- Psychotherapies and Psychological Support for Individuals Facing Psychological Distress during the COVID-19 Pandemic: A Scoping Review 94%
- The Benefits and Harms of Open Notes in Mental Health: A Delphi Survey of International Experts 94%
- Artificial Intelligence for Contextual Well-being: Protocol for an Exploratory Sequential Mixed Methods Study with Medical Students as a Social Microcosm 93%
Similar papers in this journal
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 95%
- Artificial Intelligence (AI)-based Chatbots in Promoting Health Behavioral Changes: A Systematic Review 94%
- Remote working in mental health services: a rapid umbrella review of pre-COVID-19 literature 94%
Similar papers in this journal
- Development and Evaluation of MADDIE: Method to Acquire Delivery Date Information from Electronic Health Records 92%
- Performance of Advanced Large Language Models (GPT-4o, GPT-4, Gemini 1.5 Pro, Claude 3 Opus) on Japanese Medical Licensing Examination: A Comparative Study 91%
- Digital applications to support self-management of multimorbidity: A scoping review 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.