Benchmarking Open-Source Vision-Language Models for Brain Metastasis Assessment on Single-Slice Contrast-Enhanced MRI
Kim, J.; Kim, B.-s.; Ko, J. S.; Dong, J.; Youn, S. Y.; Jang, J.; Ahn, K.-J.
Show abstract
Purpose Open-source vision-language models (VLMs) can be locally deployed without external internet access, potentially enhancing data security. This study compared the diagnostic performance of general-purpose and medical-purpose open-source VLMs and evaluated their ability to characterize brain metastases on contrast-enhanced (CE) MRI. Materials and Methods Sixty lesion-positive axial CE T1-weighted images and sixty matched lesion-negative images from 60 patients were analyzed using three general-purpose VLMs-InternVL3-8B, Qwen2.5-VL-7B-Instruct, and MiniCPM-V-4.5-and three medical-purpose VLMs-MedGemma-4B-it, LLaVA-Med v1.5, and HuatuoGPT-Vision-7B. Lesion detection performance was assessed using sensitivity, specificity, and balanced accuracy. On lesion-positive images, accuracy was evaluated for lesion count, laterality, anatomic location, enhancement pattern, necrosis, vasogenic edema, and mass effect. Model differences were assessed using Cochran's Q tests followed by pairwise McNemar tests with Benjamini-Hochberg correction. Results The median age of the study patients was 67 years (IQR, 61.0-70.5 years), and 35 patients were male (58.3%). MiniCPM-V-4.5 showed the most balanced diagnostic performance, with a sensitivity of 78.3% (95% CI, 66.4-86.9%) and a specificity of 85.0% (95% CI, 73.9-91.9%), and significantly higher balanced accuracy than all other models. Significant overall differences were observed for lesion count, laterality, location, enhancement pattern, necrosis, and mass effect, but not for vasogenic edema (FDR-adjusted P = 0.056). HuatuoGPT-Vision-7B and MedGemma-4B-it showed relatively consistent accuracy across multiple image assessment tasks, although their performance remained modest. Conclusion Our study demonstrated substantial heterogeneity in the performance of open-source VLMs in brain metastasis evaluation, and medical-purpose VLMs did not outperform general-purpose VLMs.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 96%
- From Community Acquired Pneumonia to COVID-19: A Deep Learning Based Method for Quantitative Analysis of COVID-19 on thick-section CT Scans 91%
- Impact of Non-Contrast Enhanced Imaging Input Sequences on the Generation of Virtual Contrast-Enhanced Breast MRI Scans using Neural Networks 91%
Similar papers in this journal
- Deep neural networks allow expert-level brain meningioma detection, segmentation and improvement of current clinical practice 95%
- Developing a Fully Automated Imaging Biomarker for HCC Risk Assessment via MRI-Based Tumor Segmentation and EPM 94%
- An AI-based segmentation and analysis pipeline for high-field MR monitoring of cerebral organoids 93%
Similar papers in this journal
- A Clinical Neuroimaging Platform for Rapid, Automated Lesion Detection and Personalized Post-Stroke Outcome Prediction 93%
- Foundation versus Domain-Specific Models for Cardiac Ultrasound Segmentation 89%
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 89%
Similar papers in this journal
- pyKNEEr: An image analysis workflow for open and reproducible research on femoral knee cartilage 93%
- Freewater EstimatoR using iNtErpolated iniTialization (FERNET): Toward Accurate Estimation of Free Water in Peritumoral Region Using Single-Shell Diffusion MRI Data 92%
- Image-localized Biopsy Mapping of Brain Tumor Heterogeneity: A Single-Center Study Protocol 92%
Similar papers in this journal
- Automated Tumor Segmentation and Brain Tissue Extraction from Multiparametric MRI of Pediatric Brain Tumors: A Multi-Institutional Study 95%
- Pediatric brain tumor classification using deep learning on MR-images with age fusion 95%
- Early prognostication of overall survival for pediatric diffuse midline gliomas using MRI radiomics and machine learning 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.