The Promise and Peril of Large Language Models in Digital Health: GPT-4 Personalizes Cardiovascular Patient Education but Amplifies Gender Biases
Khan, S.; Vaidean, G.
Show abstract
BackgroundGender-neutral patient education materials often overlook critical sex-based differences in cardiovascular disease (CVD). Large Language Models (LLMs) like GPT-4 offer a potential tool for personalizing health communication, but their ability to correct gender gaps without introducing new biases is unknown. MethodsWe identified seven publicly available English-language CVD prevention handouts from major health organizations. Using GPT-4 API in August 2025, we generated gender-specific revisions for a 55-year-old male and female audience via standardized prompts. We provided structured prompts instructing the model to include evidence-based, sex-specific risk factors and symptoms. Original and revised materials were evaluated using Flesch-Kincaid Reading Ease, a novel 10-point gender-inclusivity checklist, and qualitative thematic analysis. FindingsGPT-4 revisions substantially improved gender-inclusivity scores (Original median: 3.0/10, IQR=0.0-3.0; Male-tailored median: 8.0, IQR=7.0-9.0; Female-tailored median: 10.0, IQR=10.0-10.0). Readability was maintained. However, qualitative analysis revealed that while female-tailored versions excelled at incorporating biological facts (e.g., menopause), male-tailored versions often missed key clinical factors. For instance, 4 of 7 revisions failed to mention erectile dysfunction as a CVD risk marker. Revisions also occasionally relied on social stereotypes (e.g., "bottling up emotions"). Furthermore, both versions showed evidence of linguistic bias (framing female symptoms as atypical, thereby reinforcing the male-centric clinical paradigm) and gendered assumptions in recommended activities. ConclusionLLMs can rapidly improve gender-specificity in patient education but can also perpetuate harmful stereotypes and linguistic biases likely absorbed from their training data. Their use requires careful, critical oversight to avoid undermining and to potentially advance health equity. Author SummaryO_ST_ABSWhy was this study done?C_ST_ABSHeart disease affects men and women differently, but most patient education materials tend to be the same for everyone. We wanted to see if Artificial Intelligence, specifically large language models like GPT-4, could help rewrite these materials to be more accurate and helpful for each gender. What did the researchers do and find?We took seven heart disease prevention online handouts from major public health and clinical organizations. We used GPT-4 to create new versions specifically for men and for women, including male-specific risk markers (e.g., erectile dysfunction) and female-specific risk enhancers (e.g., menopause). We found that the original handouts were not very gender-specific. The AI-revised versions were dramatically better, especially for women, providing more relevant information without making the text harder to read. However, we also found the AI sometimes made mistakes, like using stereotypes about how men handle emotions. What do these findings mean?This means AI can be a powerful assistant for creating drafts, but its not a replacement for human expertise. To be used safely, a healthcare professional must always check the AIs work to ensure it is medically accurate and free from harmful stereotypes.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
- Defining Destigmatizing Design Guidelines for Use in Sexual Health-Related Digital Technologies: A Delphi Study 93%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 92%
Similar papers in this journal
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 92%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 92%
- Cracking the Code: A Scoping Review to Unite Disciplines in Tackling Legal Issues in Health Artificial Intelligence 91%
Similar papers in this journal
- Can co-designed educational interventions help consumers think critically about asking ChatGPT health questions? Results from a randomised-controlled trial 94%
- Digital Health Tools for the Passive Monitoring of Depression: A Systematic Review of Methods 92%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 92%
Similar papers in this journal
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 93%
- Multimodal Recruitment for an Internet-Based Pilot Study of Ovulation and Menstruation (OM) Health 93%
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 93%
Similar papers in this journal
- ChatGPT- versus human-generated answers to frequently asked questions about diabetes: a Turing test-inspired survey among employees of a Danish diabetes center 92%
- The Pandemic Journaling Project: A new dataset of first-person accounts of the COVID-19 pandemic 92%
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.