Red Teaming Large Language Models in Medicine: Real-World Insights on Model Behavior
Chang, C. T.-T.; Farah, H.; Gui, H.; Rezaei, S. J.; Bou-Khalil, C.; Park, Y.-J.; Swaminathan, A.; Omiye, J. A.; Kolluri, A.; Chaurasia, A.; Lozano, A.; Heiman, A.; Jia, A. S.; Kaushal, A.; Jia, A.; Iacovelli, A.; Yang, A.; Salles, A.; Singhal, A.; Narasimhan, B.; Belai, B.; Jacobson, B. H.; Li, B.; Poe, C. H.; Sanghera, C.; Zheng, C.; Messer, C.; Kettud, D. V.; Pandya, D.; Kaur, D.; Hla, D.; Dindoust, D.; Moehrle, D.; Ross, D.; Chou, E.; Lin, E.; Haredasht, F. N.; Cheng, G.; Gao, I.; Chang, J.; Silberg, J.; Fries, J. A.; Xu, J.; Jamison, J.; Tamaresis, J. S.; Chen, J. H.; Lazaro, J.; Banda, J.
Show abstract
0.BackgroundThe integration of large language models (LLMs) in healthcare offers immense opportunity to streamline healthcare tasks, but also carries risks such as response accuracy and bias perpetration. To address this, we conducted a red-teaming exercise to assess LLMs in healthcare and developed a dataset of clinically relevant scenarios for future teams to use. MethodsWe convened 80 multi-disciplinary experts to evaluate the performance of popular LLMs across multiple medical scenarios. Teams composed of clinicians, medical and engineering students, and technical professionals stress-tested LLMs with real world clinical use cases. Teams were given a framework comprising four categories to analyze for inappropriate responses: Safety, Privacy, Hallucinations, and Bias. Prompts were tested on GPT-3.5, GPT-4.0, and GPT-4.0 with the Internet. Six medically trained reviewers subsequently reanalyzed the prompt-response pairs, with dual reviewers for each prompt and a third to resolve discrepancies. This process allowed for the accurate identification and categorization of inappropriate or inaccurate content within the responses. ResultsThere were a total of 382 unique prompts, with 1146 total responses across three iterations of ChatGPT (GPT-3.5, GPT-4.0, GPT-4.0 with Internet). 19.8% of the responses were labeled as inappropriate, with GPT-3.5 accounting for the highest percentage at 25.7% while GPT-4.0 and GPT-4.0 with internet performing comparably at 16.2% and 17.5% respectively. Interestingly, 11.8% of responses were deemed appropriate with GPT-3.5 but inappropriate in updated models, highlighting the ongoing need to evaluate evolving LLMs. ConclusionThe red-teaming exercise underscored the benefits of interdisciplinary efforts, as this collaborative model fosters a deeper understanding of the potential limitations of LLMs in healthcare and sets a precedent for future red teaming events in the field. Additionally, we present all prompts and outputs as a benchmark for future LLM model evaluations. 1-2 Sentence DescriptionAs a proof-of-concept, we convened an interactive "red teaming" workshop in which medical and technical professionals stress-tested popular large language models (LLMs) through publicly available user interfaces on clinically relevant scenarios. Results demonstrate a significant proportion of inappropriate responses across GPT-3.5, GPT-4.0, and GPT-4.0 with Internet (25.7%, 16.2%, and 17.5%, respectively) and illustrate the valuable role that non-technical clinicians can play in evaluating models.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 94%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 92%
- Cracking the Code: A Scoping Review to Unite Disciplines in Tackling Legal Issues in Health Artificial Intelligence 92%
Similar papers in this journal
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 95%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
- Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.