A pen mark is all you need - Incidental prompt injection attacks on Vision Language Models in real-life histopathology
Clusmann, J.; Schulz, S. J. K.; Ferber, D.; Wiest, I. C.; Fernandez, A.; Eckstein, M.; Lange, F.; Reitsam, N. G.; Kellers, F.; Schmitt, M.; Neidlinger, P.; Koop, P.-H.; Schneider, C. V.; Truhn, D.; Roth, W.; Jesinghaus, M.; Kather, J. N.; Foersch, S.
Show abstract
Vision-language models (VLMs) can analyze multimodal medical data. However, a significant weakness of VLMs, as we have recently described, is their susceptibility to prompt injection attacks. Here, the model receives conflicting instructions, leading to potentially harmful outputs. In this study, we hypothesized that handwritten labels and watermarks on pathological images could act as inadvertent prompt injections, influencing decision-making in histopathology. We conducted a quantitative study with a total of N = 3888 observations on the state-of-the-art VLMs Claude 3 Opus, Claude 3.5 Sonnet and GPT-4o. We designed various real-world inspired scenarios in which we show that VLMs rely entirely on (false) labels and watermarks if presented with those next to the tissue. All models reached almost perfect accuracies (90 - 100 %) for ground-truth leaking labels and abysmal accuracies (0 - 10 %) for misleading watermarks, despite baseline accuracies between 30-65 % for various multiclass problems. Overall, all VLMs accepted human-provided labels as infallible, even when those inputs contained obvious errors. Furthermore, these effects could not be mitigated by prompt engineering. It is therefore imperative to consider the presence of labels or other influencing features during future evaluation of VLMs in medicine and other fields.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Interpretable multimodal deep learning for real-time pan-tissue pan-disease pathology search on social media 96%
- Attention-based whole-slide image compression achieves pathologist-level pre-screening of multi-organ routine histopathology biopsies 94%
- Genomic Characterization of Lung Cancer in Never-Smokers Using Deep Learning 94%
Similar papers in this journal
- Development and validation of AI-based pre-screening of large bowel biopsies 94%
- Novel deep learning algorithm predicts the status of molecular pathways and key mutations in colorectal cancer from routine histology images 93%
- CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research 92%
Similar papers in this journal
Similar papers in this journal
- Weakly-Supervised Tumor Purity Prediction FromFrozen H&E Stained Slides 96%
- Integrative deep learning analysis improves colon adenocarcinoma patient stratification at risk for mortality 95%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.