Prompt injection attacks on vision-language models for surgical decision support
Zhang, Z.; Qadir, M. I.; Carstens, M.; Zhang, E. H.; Loiselle, M. S.; Martinus, F. M.; Mroczkowski, M. K.; Clusmann, J.; Kather, J. N.; Kolbinger, F. R.
Show abstract
ImportanceArtificial Intelligence-driven analysis of laparoscopic video holds potential to increase the safety and precision of minimally invasive surgery. Vision-language models are particularly promising for video-based surgical decision support due to their capabilities to comprehend complex temporospatial (video) data. However, the same multimodal interfaces that enable such capabilities also introduce new vulnerabilities to manipulations through embedded deceptive text or images (prompt injection attacks). ObjectiveTo systematically evaluate how susceptible state-of-the-art video-capable vision-language models are to textual and visual prompt injection attacks in the context of clinically relevant surgical decision support tasks. Design, Setting, and ParticipantsIn this observational study, we systematically evaluated four state-of-the-art vision-language models, Gemini 1.5 Pro, Gemini 2.5 Pro, GPT-o4-mini-high, and Qwen 2.5-VL, across eleven surgical decision support tasks: detection of bleeding events, foreign objects, image distortions, critical view of safety assessment, and surgical skill assessment. Prompt injection scenarios involved misleading textual prompts and visual perturbations, displayed as white text overlay, applied at varying durations. Main Outcomes and MeasuresThe primary measure was model accuracy, contrasted between baseline performance and each prompt injection condition. ResultsAll vision-language models demonstrated good baseline accuracy, with Gemini 2.5 Pro generally achieving the highest mean [standard deviation] accuracy across all tasks (0.82 [0.01]), compared to Gemini 1.5 Pro (0.70 [0.03]) and GPT-o4 mini-high (0.67 [0.06]). Across tasks, Qwen 2.5-VL censored most outputs and achieved an accuracy of (0.58 [0.03]) on non-censored outputs. Textual and temporally-varying visual prompt injections reduced the accuracy for all models. Prolonged visual prompt injections were generally more harmful than single-frame injections. Gemini 2.5 Pro showed the greatest robustness and maintained stable performance for several tasks despite prompt injections, whereas GPT-o4-mini-high exhibited the highest vulnerability, with mean (standard deviation) accuracy across all tasks declining from 0.67 (0.06) at baseline to 0.24 (0.04) under full-duration visual prompt injection (P < .001). Conclusion and RelevanceThese findings indicate the critical need for robust temporal reasoning capabilities and specialized guardrails before vision-language models can be safely deployed for real-time surgical decision support. Key PointsO_ST_ABSQuestionC_ST_ABSAre video vision-language models (VLMs) susceptible to textual and visual prompt injection attacks when used for surgical decision support tasks? FindingTextual and visual prompt injection attacks consistently degraded the performance of four state-of-the-art VLMs across eleven surgical tasks. Gemini 2.5 Pro was most robust to textual and visual prompt injection attacks, whereas GPT-o4-mini-high was most vulnerable. Prolonged visual injections had a greater negative impact than single-frame injections. MeaningPresent-generation video VLMs are highly vulnerable to textual and visual prompt injection attacks. This critical safety vulnerability must be addressed before their integration into surgical decision support systems.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 94%
- Enhancing Privacy-Preserving Deployable Large Language Models for Perioperative Complication Detection: A Targeted Strategy with LoRA Fine-tuning 94%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 94%
Similar papers in this journal
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 94%
- An Open-Source Tool for Automated Human-Level Circling Behavior Detection 93%
- Expert Surgeons and Deep Learning Models Can Predict the Outcome of Surgical Hemorrhage from One Minute of Video 93%
Similar papers in this journal
- Benchmarking transformer-based models for medical record deidentification: A single centre, multi-specialty evaluation 94%
- EYE-Llama, an in-domain large language model for ophthalmology 93%
- A multi-task domain-adapted model to predict chemotherapy response from mutations in recurrently altered cancer genes 91%
Similar papers in this journal
- Annotation-free multi-organ anomaly detection in abdominal CT using free-text radiology reports: A multi-center retrospective study 92%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 92%
- Deep Learning Prediction of Biomarkers from Echocardiogram Videos 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.