A Framework for Evaluating the Efficacy of Foundation Embedding Models in Healthcare
Xu, S.; Gui, H.; Rotemberg, V.; Wang, T.; Chen, Y.; Daneshjou, R.
Show abstract
Recent interest has surged in building large-scale foundation models for medical applications. In this paper, we propose a general framework for evaluating the efficacy of these foundation models in medicine, suggesting that they should be assessed across three dimensions: general performance, bias/fairness, and the influence of confounders. Utilizing Googles recently released dermatology embedding model and lesion diagnostics as examples, we demonstrate that: 1) dermatology foundation models surpass state-of-the-art classification accuracy; 2) general-purpose CLIP models encode features informative for medical applications and should be more broadly considered as a baseline; 3) skin tone is a key differentiator for performance, and the potential bias associated with it needs to be quantified, monitored, and communicated; and 4) image quality significantly impacts model performance, necessitating that evaluation results across different datasets control for this variable. Our findings provide a nuanced view of the utility and limitations of large-scale foundation models for medical AI.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An Inherently Interpretable AI model improves Screening Speed and Accuracy for Early Diabetic Retinopathy 93%
- Assessing generalizability of an AI-based visual test for cervical cancer screening 93%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 92%
Similar papers in this journal
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 92%
- A Deep Learning Based Smartphone Application for Early Detection of Nasopharyngeal Carcinoma Using Endoscopic Images 91%
- Equity-Enhanced Glaucoma Progression Prediction from OCT with Knowledge Distillation 90%
Similar papers in this journal
- BenchXAI: Comprehensive Benchmarking of Post-hoc Explainable AI Methods on Multi-Modal Biomedical Data 93%
- The mathematics of erythema: Development of machine learning models for artificial intelligence assisted measurement and severity scoring of radiation induced dermatitis 93%
- Fusion of Electronic Health Records and Radiographic Images for a Multimodal Deep Learning Prediction Model of Atypical Femur Fractures 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.