Back

Performance and Adversarial Vulnerability of Vision Language Models in Computer Tomography

Choi, B.; Hong, S.; Park, M. W.; Jun, T. J.; Suh, J.

2025-10-13 health informatics
10.1101/2025.10.10.25337723 medRxiv
Show abstract

This study investigates the performance and vulnerability of Vision Language Models (VLMs) in interpreting computed tomography (CT). In a factorial experiment, four leading VLMs (Google, OpenAI, Anthropic, Alibaba) was used to classified 240 kidney CT scans into tumor, cyst, or normal categories under three prompt conditions. A Neutral prompt requested simple interpretation, while adversarial Benign prompts aimed to mislead, and Pressure prompts simulated clinical overload. Key performance metrics, including accuracy, precision, and recall, were evaluated. The overall 3-category classification accuracy under neutral conditions was 48.4%, with Googles VLM achieving the highest individual accuracy (60.0%), followed by OpenAI (49.2%), Alibaba (42.5%), and Anthropic (42.1%). The introduction of adversarial prompts significantly degraded performance, with overall accuracy decreasing to 39.2% (p<0.001) under benign prompts and 43.3% (p=0.024) under pressure prompts. These prompts also induced significant prediction skews; for instance, pressure prompts systematically biased all models toward normal classifications. In conclusion, current VLMs demonstrated modest accuracy for kidney CT classification and were highly vulnerable to adversarial manipulation. These findings raise critical concerns about their reliability and highlight the urgent need for extensive validation before any clinical implementation.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.