Response consistency of ChatGPT-4o for Type 2 Diabetes Nutrition and Physical-activity Recommendations: A Pilot NLP-based Assessment of GPT outputs
Zhang, Y.; Liu, X.-J.; Hu, Q.; Galaviz, K. I.; Casanova, I. G.; Colditz, J.; Valdez, D.
Show abstract
Generative AI tools such as ChatGPT are increasingly used by the public to seek guidance on diet and physical activity for type 2 diabetes (T2D) prevention and management. However, the consistency of model outputs across different users and disease-stage scenarios remains insufficiently characterized. This pilot study aims to evaluate the word-level and semantic-level consistency of GPT-4os diet and physical activity responses for type 2 diabetes prevention and management. We designed 12 prompts covering four categories: prediabetes, diagnosed type 2 diabetes (T2D), diagnosed T2D with complications, and general questions that did not specify dysglycemia stage. Word-level similarity was quantified with Term Frequency-Inverse Document Frequency (TF-IDF) cosine scores; sentence-level semantic similarity was measured using large language models (LLMs) - DeBERTa-v3 MNLI to calculate the entailment probabilities. The results showed that mean cosine similarity across users was moderate (0.44-0.66), whereas mean entailment similarity was higher (0.68-0.81). Across stages, word-level similarity was low to moderate (0.28-0.63) and entailment similarity remained moderate to high (0.63-0.80). Low similarity commonly referenced distinct food choices, operational details, safety warnings, and stage-specific suggestions. GPT-4o generated semantically consistent but variably detailed responses and the moderate semantic variation suggested limited differentiation of response content across diabetes-related stages in this pilot consistency assessment.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enhancing Early Detection of Cognitive Decline in the Elderly through Ensemble of NLP Techniques: A Comparative Study Utilizing Large Language Models in Clinical Notes 92%
- irAE-GPT: Leveraging large language models to identify immune-related adverse events in electronic health records and clinical trial datasets 91%
- Consistent Performance of GPT-4o in Rare Disease Diagnosis Across Nine Languages and 4967 Cases 90%
Similar papers in this journal
Similar papers in this journal
- ChatGPT- versus human-generated answers to frequently asked questions about diabetes: a Turing test-inspired survey among employees of a Danish diabetes center 94%
- Demographic and socioeconomic determinants of access to care: A subgroup disparity analysis using new equity-focused measurements 92%
- Optimization of nutritional strategies using a mechanistic computational model in prediabetes: Application to the J-DOIT1 study data 92%
Similar papers in this journal
- Artificial Intelligence's Contribution to Biomedical Literature Search: Revolutionizing or Complicating? 92%
- Racial disparities in continuous glucose monitoring-based 60-min glucose predictions among people with type 1 diabetes 92%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.