Prompt Engineering in Large Language Models for Patient Education: A Systematic Review
Mudrik, A.; Nadkarni, G.; Efros, O.; Soffer, S.; Klang, E.
Show abstract
BackgroundLarge language models (LLMs) have shown promise in generating patient-friendly medical content, but their outputs often vary in accuracy, readability, and relevance. Prompt engineering--structuring inputs to guide LLM responses--may improve the quality of educational materials, yet its impact on patient education remains unclear. ObjectivesTo systematically review whether prompt engineering improves readability, accuracy, and usability of LLM-generated content for patient education. MethodsWe conducted a systematic review in accordance with PRISMA guidelines. PubMed, Scopus, and Web of Science were searched for original studies evaluating prompt engineering techniques in patient education. Data were extracted on prompt types, LLM models used, and outcomes. Risk of bias was assessed using the QUADAS-2 tool, and a narrative synthesis was performed. ResultsOur search identified five studies that met our criteria, focusing on answering patient questions and generating medical information. Prompt engineering techniques included instruction-based, elaborated, role-defining, scene-defining, and domain-specific prompts. Structured prompting improved accuracy and comprehensiveness in several cases, particularly when specific formats or custom instructions were used. Readability gains were notable when prompts explicitly requested simpler language and reading levels, though some strategies unintentionally increased complexity. Variability in effectiveness across LLMs and prompt types was observed. ConclusionPrompt engineering can enhance the clarity and, in some cases, the accuracy of LLM-generated patient education materials. However, benefits vary by model and strategy. Standardized approaches and further research are needed to optimize prompts, minimize bias, and support reliable, accessible patient communication.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 94%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 93%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 93%
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 95%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 93%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 92%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 91%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 90%
Similar papers in this journal
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 92%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 91%
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.