Multimodal Foundation Models for Medical Imaging - A Systematic Review and Implementation Guidelines
Huang, S.-C.; Jensen, M. E. K.; Yeung-Levy, S.; Lungren, M. P.; Poon, H.; Chaudhari, A.
Show abstract
Advancements in artificial intelligence (AI) offer promising solutions for enhancing clinical workflows and patient care, potentially revolutionizing healthcare delivery. However, the traditional paradigm of AI integration in healthcare is limited by models that rely on single input modalities during training and require extensive labeled data, failing to capture the multimodal nature of medical practice. Multimodal foundation models, particularly Large Vision Language Models (VLMs), have the potential to overcome these limitations by processing diverse data types and learning from large-scale unlabeled datasets or natural pairs of different modalities, thereby significantly contributing to the development of more robust and versatile AI systems in healthcare. In this review, we establish a unified terminology for multimodal foundation models for medical imaging applications and provide a systematic analysis of papers published between 2012 and 2024. In total, we screened 1,144 papers from medical and AI domains and extracted data from 97 included studies. Our comprehensive effort aggregates the collective knowledge of prior work, evaluates the current state of multimodal AI in healthcare, and delineates both prevailing limitations and potential growth areas. We provide implementation guidelines and actionable recommendations for various stakeholders, including model developers, clinicians, policymakers, and dataset curators.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 94%
- Enhancing Fairness in Disease Prediction by Optimizing Multiple Domain Adversarial Networks 94%
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 94%
Similar papers in this journal
- ARA: accurate, reliable and active histopathological image classification framework with Bayesian deep learning 95%
- Dual Adversarial Deconfounding Autoencoder for joint batch-effects removal from multi-center and multi-scanner radiomics data 94%
- Generating synthetic data in digital pathology through diffusion models: a multifaceted approach to evaluation 94%
Similar papers in this journal
- Using Adversarial Images to Assess the Stability of Deep Learning Models Trained on Diagnostic Images in Oncology 94%
- Actionability of Synthetic Data in a Heterogeneous and Rare Healthcare Demographic; Adolescents and Young Adults (AYAs) with Cancer 93%
- Histology-based Prediction of Therapy Response to Neoadjuvant Chemotherapy for Esophageal and Esophagogastric Junction Adenocarcinomas Using Deep Learning 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.