Evaluating the Multimodal Capabilities of Generative AI in Complex Clinical Diagnostics
Schubert, M. C.; Lasotta, M.; Sahm, F.; Wick, W.; Venkataramani, V.
Show abstract
In the rapidly evolving landscape of artificial intelligence (AI) in healthcare, the study explores the diagnostic capabilities of Generative Pre-trained Transformer 4 Vision (GPT-4V) in complex clinical scenarios involving both medical imaging and textual patient data. Conducted over a week in October 2023, the study employed 93 cases from the New England Journal of Medicines image challenges. These cases were categorized into four types based on the nature of the imaging data, ranging from radiological scans to pathological slides. GPT-4Vs diagnostic performance was evaluated using multimodal inputs (text and image), text-only, and image-only prompts. The results indicate that GPT-4Vs diagnostic accuracy was highest when provided with multimodal inputs, aligning with the confirmed diagnoses in 80.6% of cases. In contrast, text-only and image-only inputs yielded lower accuracies of 66.7% and 45.2%, respectively (after correcting for random guessing: multimodal: 70.5 %, text only: 54.3 %, image only: 29.3 %). No significant variation was observed in the models performance across different types of images or medical specialties. The study substantiates the utility of multimodal AI models like GPT-4V as potential aids in clinical diagnostics. However, the proprietary nature of the models training data and architecture warrants further investigation to uncover biases and limitations. Future research should aim to corroborate these findings with real-world clinical data while considering ethical and privacy concerns.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 92%
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 92%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 91%
Similar papers in this journal
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 92%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 92%
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 92%
Similar papers in this journal
- GenECG: A synthetic image-based ECG dataset to augment artificial intelligence-enhanced algorithm development 89%
- ChatGPT in glioma patient adjuvant therapy decision making: ready to assume the role of a doctor in the tumour board? 87%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 87%
Similar papers in this journal
- Data-driven Discovery of Mathematical and Physical Relations in Oncology Data using Human-understandable Machine Learning 90%
- Predicting Ward Transfer Mortality with Machine Learning 89%
- Predicting the disease outcome in COVID-19 positive patients through Machine Learning: a retrospective cohort study with Brazilian data 88%
Similar papers in this journal
- Refining LLMs Outputs with Iterative Consensus Ensemble (ICE) 92%
- Fusion of Electronic Health Records and Radiographic Images for a Multimodal Deep Learning Prediction Model of Atypical Femur Fractures 92%
- BenchXAI: Comprehensive Benchmarking of Post-hoc Explainable AI Methods on Multi-Modal Biomedical Data 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.