Seeing Nothing, Saying Something: The Lack of Visual Grounding and Confabulation in Gemini Models for Histopathology
Hasan, M. M.; Tozal, M. E.; Ayhan, M. S.
Show abstract
Large vision-language models (VLMs) have demonstrated remarkable perfor- mance on computational pathology benchmarks, yet their reliability under adversarial or vacuous inputs remains poorly understood. This paper examines the visual grounding behaviour of two Gemini models Gemini 3.0 Flash Pre- view (gemini-flash) and Gemini 3.1 Pro Preview (gemini-pro) on a well known histopathology classification task, and probes for confabulation using a adver- sarial blank-image set. On the real histopathology dataset both models achieve near-perfect accuracy (98.75% - 100%) across three temperatures (0.0, 0.5, 1.0) and three independent runs. On a controlled adversarial set of blank white images labelled as either benign or malignant, however, a stark divergence emerges. Gemini-flash consistently acknowledges the absence of visual content and assigns zero confidence, while Gemini-pro fabricates detailed, clinically plausible histo- logical descriptions and reports high confidence (mean {approx} 0.95) across the same blank inputs, a behaviour we term confident confabulation. The confabulation rate of gemini-pro reaches 77.8% image-responses at temperature 0.0, dropping to 44.4% at temperature 0.5 and rising to 66.7% at temperature 1.0, while gemini- flash records 0% at all temperatures. These findings raise important questions about the safety and trustworthiness of VLMs in clinical decision-support con- texts, and underscore the need for comprehensive evaluation beyond standard accuracy metrics.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Aggregation of Cohorts for Histopathological Diagnosis with Deep Morphological Analysis 95%
- Generating synthetic data in digital pathology through diffusion models: a multifaceted approach to evaluation 95%
- ARA: accurate, reliable and active histopathological image classification framework with Bayesian deep learning 94%
Similar papers in this journal
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 95%
- Assessing generalizability of an AI-based visual test for cervical cancer screening 95%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 93%
Similar papers in this journal
Similar papers in this journal
- Using Adversarial Images to Assess the Stability of Deep Learning Models Trained on Diagnostic Images in Oncology 94%
- Actionability of Synthetic Data in a Heterogeneous and Rare Healthcare Demographic; Adolescents and Young Adults (AYAs) with Cancer 93%
- frenchFISH: Poisson models for quantifying DNA copy-number from fluorescence in situ hybridisation of tissue sections 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.