Performance and Adversarial Vulnerability of Vision Language Models in Computer Tomography
Choi, B.; Hong, S.; Park, M. W.; Jun, T. J.; Suh, J.
Show abstract
This study investigates the performance and vulnerability of Vision Language Models (VLMs) in interpreting computed tomography (CT). In a factorial experiment, four leading VLMs (Google, OpenAI, Anthropic, Alibaba) was used to classified 240 kidney CT scans into tumor, cyst, or normal categories under three prompt conditions. A Neutral prompt requested simple interpretation, while adversarial Benign prompts aimed to mislead, and Pressure prompts simulated clinical overload. Key performance metrics, including accuracy, precision, and recall, were evaluated. The overall 3-category classification accuracy under neutral conditions was 48.4%, with Googles VLM achieving the highest individual accuracy (60.0%), followed by OpenAI (49.2%), Alibaba (42.5%), and Anthropic (42.1%). The introduction of adversarial prompts significantly degraded performance, with overall accuracy decreasing to 39.2% (p<0.001) under benign prompts and 43.3% (p=0.024) under pressure prompts. These prompts also induced significant prediction skews; for instance, pressure prompts systematically biased all models toward normal classifications. In conclusion, current VLMs demonstrated modest accuracy for kidney CT classification and were highly vulnerable to adversarial manipulation. These findings raise critical concerns about their reliability and highlight the urgent need for extensive validation before any clinical implementation.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 95%
- Assessing generalizability of an AI-based visual test for cervical cancer screening 95%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 95%
Similar papers in this journal
- On evaluation metrics for medical applications of artificial intelligence 95%
- Generating synthetic data in digital pathology through diffusion models: a multifaceted approach to evaluation 94%
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 94%
Similar papers in this journal
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 95%
- A Deep Learning Based Smartphone Application for Early Detection of Nasopharyngeal Carcinoma Using Endoscopic Images 93%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 93%
Similar papers in this journal
Similar papers in this journal
- Using Adversarial Images to Assess the Stability of Deep Learning Models Trained on Diagnostic Images in Oncology 94%
- DeepPhe-CR: Natural Language Processing Software Services for Cancer Registrar Case Abstraction 93%
- Towards Predicting 30-Day Readmission among Oncology Patients: Identifying Timely and Actionable Risk Factors 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.