Evaluation of Large Language Models for Post-Cystectomy Sexual Health Counseling in Women: A Pilot Study
Shafau, F.; Dave, A. A.; Omole, I.; Guzman, T.; Rehman, N.; Enemchukwu, E.; Bresler, L.
Show abstract
Abstract Objective To evaluate the adherence to guidelines and readability of large language model-generated sexual health information related to female sexual dysfunction following cystectomy, and to determine whether adherence differs across models and prompt formats. A secondary objective was to introduce an analytic strategy using principal component analysis to examine the dimensions of readability metrics. Methods Three large language models (LLMs), ChatGPT, Gemini, and Perplexity were prompted with six clinical questions related to sexual function after cystectomy. Questions were phrased in long-form and short-form language. Responses were independently graded by two reviewers, derived from guideline recommendations. Linear mixed-effects models predicted adherence as functions of LLM, prompt, and reviewer, with clinical questions as a random intercept. Readability was assessed using five metrics, and principal component analysis (PCA) was used to determine latent structure. Results ChatGPT demonstrated the highest (estimated marginal mean [emm] = 0.769), outperforming Gemini (0.499) and Perplexity (0.457). Shorter, less complex prompts elicited higher adherence than more complex, clinical prompts. All models produced content that exceeded recommended reading levels. PCA demonstrated that a single dominant component accounted for 76.7% of variance across readability indices, indicating a shared underlying construct. Conclusion ChatGPT produced the most guideline-concordant information overall. High linguistic complexity was seen across models, highlighting a barrier to patient comprehension. These findings characterize large language models as variable medical information systems whose outputs rely heavily on prompt structure and model type.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 94%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 92%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 92%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 92%
- Factors Influencing Precision Medicine Knowledge and Attitudes 92%
- ChatGPT- versus human-generated answers to frequently asked questions about diabetes: a Turing test-inspired survey among employees of a Danish diabetes center 92%
Similar papers in this journal
- The effect of digital-enabled multidisciplinary therapy conferences on efficiency and quality of the decision making in prostate-cancer care 93%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 91%
- Impact of the Federated Data Platform's digital surgery scheduling system on elective theatre utilisation at an NHS Trust: an interrupted time series analysis 90%
Similar papers in this journal
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 93%
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 93%
- Multimodal Recruitment for an Internet-Based Pilot Study of Ovulation and Menstruation (OM) Health 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.