A Comprehensive Study of GPT-4V's Multimodal Capabilities in Medical Imaging
Li, Y.; Liu, Y.; Wang, Z.; Liang, X.; Liu, L.; Wang, L.; Cui, L.; Tu, Z.; Wang, L.; Zhou, L.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWThis paper presents a comprehensive evaluation of GPT-4Vs capabilities across diverse medical imaging tasks, including Radiology Report Generation, Medical Visual Question Answering (VQA), and Visual Grounding. While prior efforts have explored GPT-4Vs performance in medical imaging, to the best of our knowledge, our study represents the first quantitative evaluation on publicly available benchmarks. Our findings highlight GPT-4Vs potential in generating descriptive reports for chest X-ray images, particularly when guided by well-structured prompts. However, its performance on the MIMIC-CXR dataset benchmark reveals areas for improvement in certain evaluation metrics, such as CIDEr. In the domain of Medical VQA, GPT-4V demonstrates proficiency in distinguishing between question types but falls short of prevailing benchmarks in terms of accuracy. Furthermore, our analysis finds the limitations of conventional evaluation metrics like the BLEU score, advocating for the development of more semantically robust assessment methods. In the field of Visual Grounding, GPT-4V exhibits preliminary promise in recognizing bounding boxes, but its precision is lacking, especially in identifying specific medical organs and signs. Our evaluation underscores the significant potential of GPT-4V in the medical imaging domain, while also emphasizing the need for targeted refinements to fully unlock its capabilities.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BenchXAI: Comprehensive Benchmarking of Post-hoc Explainable AI Methods on Multi-Modal Biomedical Data 93%
- Fusion of Electronic Health Records and Radiographic Images for a Multimodal Deep Learning Prediction Model of Atypical Femur Fractures 93%
- ISMI-VAE: A Deep Learning Model for Classifying Disease Cells Using Gene Expression and SNV Data 92%
Similar papers in this journal
Similar papers in this journal
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 95%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 94%
- Classification of Hyper-scale Multimodal Imaging Datasets 93%
Similar papers in this journal
Similar papers in this journal
- SN-FPN: Self-attention Nested Feature Pyramid Network for Digital Pathology Image Segmentation 93%
- An Accurate and Explainable Deep Learning System Improves Interobserver Agreement in the Interpretation of Chest Radiograph 92%
- Bayesian automatic screening of pneumoniaand lung lesions localization from CT scans. Acombined method toward a more user-centredand explainable approach 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.