Leveraging large language models to address common vaccination myths and misconceptions
Reis, F.; Bayer, L. J.; Malerczyk, C.; Lenz, C.; von Eiff, C.
Show abstract
Large language models (LLMs) are increasingly used by the public to seek health information, yet their reliability in addressing common vaccine myths remains unclear. We conducted an exploratory multi-vendor evaluation of three LLMs (GPT-5, Gemini 2.5 Flash, Claude Sonnet 4) using officially curated vaccination myths from Germanys public health institution and two realistic user framings as prompts: a curious skeptic and a convinced believer. All model responses were independently evaluated by two blinded medical experts for misconception addressal (binary), scientific accuracy, and communication clarity (5-point Likert scales). Additionally, blinded marketing experts ranked models for lay communication clarity, and Flesch-Kincaid Reading Ease scores were computed for all outputs. Across all myths, prompts, and models (11 x 2 x 3 = 66 rating items), medical raters found 100% successful refutation of misinformation. Scientific accuracy and clarity ratings were high and tightly clustered (median 4.0-4.5), with no combined score below 3 and substantial inter-rater agreement. Marketing experts independently ranked Gemini 2.5 Flash and GPT-5 highest for lay clarity, with Claude Sonnet 4 consistently less favored. Readability analysis revealed generally low accessibility, particularly for the convinced believer framing and for Claude Sonnet 4 outputs. Our findings suggest that current general-purpose LLMs can deliver accurate debunking of widely documented vaccine myths under realistic conditions, but that linguistic complexity and framing-sensitive style may limit accessibility. Careful integration of LLMs into public health channels, alongside transparent sourcing and readability optimization, could enable these models to be used as scalable tools for debunking vaccine myths.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Mild Adverse Events of Sputnik V Vaccine Extracted from Russian Language Telegram Posts via BERT Deep Learning Model 93%
- Users’ Reactions on Announced Vaccines against COVID-19 Before Marketing in France: Analysis of Twitter posts 91%
- Has the pandemic enhanced and sustained digital health-seeking behaviour? A big data interrupted time-series analysis of Google Trends 91%
Similar papers in this journal
Similar papers in this journal
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 91%
- Simulated Misuse of Large Language Models and Clinical Credit Systems 91%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 90%
Similar papers in this journal
- Factors indicating intention to vaccinate with a COVID-19 vaccine among older U.S. Adults 93%
- A randomized controlled trial of a video intervention shows evidence of increasing COVID-19 vaccination intention 92%
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.