Multimodal Large Language Models are Generalist Medical Image Interpreters
Han, T.; Adams, L. C.; Nebelung, S.; Kather, J. N.; Bressem, K. K.; Truhn, D.
Show abstract
Medicine is undergoing a transformation with the integration of Artificial Intelligence (AI). Traditional AI models, though clinically useful and often matching or surpassing expert clinicians in specific tasks, face a scalability challenge due to the necessity of developing individual models for each task. Therefore, there is a push towards foundation models that are applicable to a wider set of tasks. Our study showcases how non-domain-specific, publicly available vision-language models can be employed as general foundation models for medical applications. We test our paradigm across four medical disciplines - pathology, dermatology, ophthalmology, and radiology - focusing on two use-cases within each discipline. We find that our approach beats existing pre-training methods and is competitive to domain-specific foundation models that require vast amounts of domain-specific training images. We also find that large vision-language models are data efficient and do not require large annotated datasets to reach competitive performance. This allows for the development of new or improved AI models in areas of medicine where data is scarce and will accelerate medical progress towards true multimodal foundation models.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Understanding the robustness of vision-language models to medical image artefacts 96%
- STPath: A Generative Foundation Model for Integrating Spatial Transcriptomics and Whole Slide Images 94%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 94%
Similar papers in this journal
- Subpopulation-specific Machine Learning Prognosis for Underrepresented Patients with Double Prioritized Bias Correction 94%
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 93%
- A user-friendly tool for cloud-based whole slide image segmentation, with examples from renal histopathology 93%
Similar papers in this journal
- Unraveling the complexity of rat object vision requires a full convolutional network - and beyond 94%
- Obtaining Spatially Resolved Tumor Purity Maps Using Deep Multiple Instance Learning In A Pan-cancer Study 93%
- Bi-level Graph Learning Unveils Prognosis-Relevant Tumor Microenvironment Patterns in Breast Multiplexed Digital Pathology 92%
Similar papers in this journal
Similar papers in this journal
- Generative AI Enables Medical Image Segmentation in Ultra Low-Data Regimes 96%
- Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms 95%
- Accurate recognition of colorectal cancer with semi-supervised deep learning on pathological images 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.