Evaluating the accuracy and consistency of ChatGPT for the management of type 2 diabetes: A cross-sectional study
Cutler, D.; Van Bakel, T.; Olar, P.; Fralick, M.
Show abstract
Large language models (LLMs) have fundamentally changed how patients and clinicians retrieve information; however, it is unclear how accurate and consistent widely available LLMs are in answering questions related to medical information. Our objective was to evaluate the accuracy and consistency of ChatGPT in answering questions related to the management of type 2 diabetes mellitus (T2DM). Three users asked ChatGPT 13 questions pertaining to medications from the top five most common classes of T2DM medications. A response was labelled inconsistent if the response provided to one user differed from the response provided to at least one other user in the same domain for the same medication. A response was labelled as inaccurate if the information provided by ChatGPT was incorrect based on the most recent FDA-approved drug label, in addition to review by an expert reviewer. Additionally, one user asked ChatGPT 26 basic questions related to the management of T2DM, in which the answer was categorized as correct or incorrect. We summarized all results using descriptive statistics. ChatGPT delivered inaccurate responses in seven out of 13 domains and inconsistent responses in seven out of 13 domains for drugs in all five classes of T2DM medication. Of ChatGPTs responses to the 26 basic T2DM treatment questions, 7 (26%) were incorrect. In this cross-sectional study, we identified that it was common for ChatGPT to provide incorrect or inconsistent responses to enquiries related to the management of type 2 diabetes.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Application of Pharmacological Statin-Associated Muscle Symptoms Phenotyping Algorithms Using Structured and Unstructured Electronic Health Records Data 92%
- Determining prescriptions in electronic health care (EHR) data: methods for development of standardised, reproducible drug codelists 92%
- Automating Evaluation of LLM-generated Responses to Patient Questions about Rare Diseases 91%
Similar papers in this journal
Similar papers in this journal
- Development and Evaluation of MADDIE: Method to Acquire Delivery Date Information from Electronic Health Records 91%
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 90%
- Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications 90%
Similar papers in this journal
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 93%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 91%
- Can co-designed educational interventions help consumers think critically about asking ChatGPT health questions? Results from a randomised-controlled trial 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.