Drug or Pokemon? An analysis of the ability of large language models to discern fabricated medications
Henry, K.; Smith, B.; Zhao, X.; Blotske, K.; Murray, B.; Gao, Y.; Smith, S. E.; Barreto, E.; Bauer, S.; Sohn, S.; Liu, T.; Bennett, T. D.; Cohen, M.; Abdulnour, R.-E. E.; Sikora, A.
Show abstract
BackgroundThe use of large language models (LLMs) is increasing in the medical field; however, LLMs are often subject to "confabulations." Notably, LLMs have vulnerability to adversarial attacks, or fabricated details within prompts, which is concerning given both health misinformation and inadvertent errors in the medical record. This purpose of this study was to determine the effect of adversarial attacks by embedding one fabricated medication into a list of existing medicines. MethodsA total of 250 cases were created, which included 4-6 medications and one fabricated medication (a Pokemon character). Four LLMs (GPT-4o-mini, Gemma-3-27B-IT, Llama-3.3-70B-Instruct, and Qwen3-32B) were tested in triplicate for both dosing information and disease indication with a default prompt, mitigation prompt, and the default prompt with a temperature of 0. If the LLM responded as if the Pokemon were a real medication, it was deemed a confabulation. The primary outcome was the rate of confabulations; exact paired-permutation tests were used to evaluate differences among LLMs and prompting approaches. ResultsConfabulation rates for the default and temperature 0 drug dosing prompt ranged from 86-98.8% and 86.9-98.8% across models, respectively, and from 42-95.6% and 41.6-95.5% for the indication prompt. Incorporating the mitigation prompt substantially reduced confabulation rates to 8.3-76.3% (dosing) and 1.7-28.3% (indication). The best-performing model, Llama-3.3-70B-Instruct, demonstrated confabulation rates spanning 1.7-91.9% (p<0.001). ConclusionsLLMs are susceptible to adversarial attacks, especially with medications. Further model improvement is imperative before LLMs are considered safe and reliable for routine use in the medical field.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 94%
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 93%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 93%
Similar papers in this journal
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 91%
- Natural language processing to evaluate texting conversations between patients and healthcare providers during COVID-19 Home-Based Care in Rwanda at scale 91%
- Evaluating Anti-LGBTQIA+ Medical Bias in Large Language Models 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.