Calm-Vlm: Calibration And Selective Prediction In Vision-Language Models For Reliable Brain MRI Classification
Dhinagar, N. J.; Jagad, C.; Senthilkumar, P.; Thomopoulos, S. I.; Khan, M. H.; Liew, S.-L.; the ENIGMA-Stroke Recovery Working Group, ; Banaj, N.; Boric, M. R.; Boyd, L. A.; Brodtmann, A.; Cassidy, J. M.; Conforto, A. B.; Cramer, S. C.; Dula, A. N.; Geranmayeh, F.; Gregory, C. M.; Hordacre, B.; Jaywant, A.; Kautz, S. A.; Leech, K. A.; Lotze, M.; Mataro, M.; Piras, F.; Rosario, E. R.; Sanossian, N.; Schambra, H. M.; Schweighofer, N.; Seo, N. J.; Soekadar, S. R.; Thielman, G. T.; Winstein, C.; Wittenberg, G. F.; Wong, K. A.; Thompson, P. M.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWRecent advances in vision-language models (VLMs) have demonstrated strong multimodal capabilities for medical image analysis. However, their confidence in diagnostic predictions is often unclear, limiting adoption in clinical settings. We introduce CALM-VLM (CALibration Mechanism for Vision-Language Models), which integrates confidence calibration and selective prediction into a generative 3D VLM. To create CALM-VLM, we fine-tuned the Med3DVLM architecture for Alzheimers disease (AD) and stroke classification as initial test cases. To improve reliability, we incorporated temperature scaling on the VLMs generative outputs. The calibrated model then selectively abstained from predictions when uncertain; this also improved its diagnostic accuracy. Experiments across multi-site MRI datasets, from 10 countries worldwide, show that CALM-VLM improved confidence relative to uncalibrated VLMs. Coverage-adjusted test receiver-operator characteristic curve-area under the curve (ROC-AUC) increased by 5% to 13% for both diagnostic tasks across independent test sets. Our calibrated VLM achieved a test ROC-AUC of 0.951 for AD classification and 0.905 for stroke classification. These findings highlight the importance of calibrated, uncertainty-aware VLMs for trustworthy neuroimaging AI.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- WMH-DualTasker: A weakly-supervised deep learning model for automated white matter hyperintensities segmentation and visual rating prediction 96%
- OpenMAP-T1: A Rapid Deep Learning Approach to Parcellate 280 Anatomical Regions to Cover the Whole Brain 95%
- Brain Age Prediction: Deep Models Need a Hand to Generalize 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Multidimensional analysis and detection of informative features in diffusion MRI measurements of human white matter 92%
- CNN MouseNet: A biologically constrained convolutional neural network model for mouse visual cortex 92%
- Searching through functional space reveals distributed visual, auditory, and semantic coding in the human brain 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.