Evaluation of Prompts to Simplify Cardiovascular Disease Information Using a Large Language Model
Mishra, V.; Sarraju, A.; Kalwani, N. M.; Dexter, J. P.
Show abstract
AI chatbots powered by large language models (LLMs) are emerging as an important source of public-facing medical information. Generative models hold promise for producing tailored guidance at scale, which could advance health literacy and mitigate well-known disparities in the accessibility of health-protective information. In this study, we highlight an important limitation of basic approaches to AI-powered text simplification: when given a zero-shot or one-shot simplification prompt, GPT-4 often responds by omitting critical details. To address this limitation, we developed a new prompting strategy, which we term rubric prompting. Rubric prompts involve a combination of a zero-shot simplification prompt with brief reminders about important topics to address. Using rubric prompts, we generate recommendations about cardiovascular disease prevention that are more complete, more readable, and have lower syntactic complexity than baseline responses produced without prompt engineering. This analysis provides a blueprint for rigorous evaluation of AI model outputs in medicine.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 90%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 90%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 89%
Similar papers in this journal
Similar papers in this journal
- Can co-designed educational interventions help consumers think critically about asking ChatGPT health questions? Results from a randomised-controlled trial 93%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 92%
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 91%
Similar papers in this journal
- Scientific hypothesis generation process in clinical research: a secondary data analytic tool versus experience study protocol 92%
- A conversational agent for providing personalized PrEP support: Protocol for chatbot implementation 89%
- Strategies and Tools for electronic health records and physician workflow alignment: A scoping review protocol 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.