Visual-Textual Integration in LLMs for Medical Diagnosis: A Quantitative Analysis
Agbareia, R.; Omar, M.; Soffer, S.; Glicksberg, B. S.; Nadkarni, G.; Klang, E.
Show abstract
Background and AimVisual data from images is essential for many medical diagnoses. This study evaluates the performance of multimodal Large Language Models (LLMs) in integrating textual and visual information for diagnostic purposes. MethodsWe tested GPT-4o and Claude Sonnet 3.5 on 120 clinical vignettes with and without accompanying images. Each vignette included patient demographics, a chief complaint, and relevant medical history. Vignettes were paired with either clinical or radiological images from two sources: 100 images from the OPENi database and 20 images from recent NEJM challenges, ensuring they were not in the LLMs training sets. Three primary care physicians served as a human benchmark. We analyzed diagnostic accuracy and the models explanations for a subset of cases. ResultsLLMs outperformed physicians in text-only scenarios (GPT-4o: 70.8%, Claude Sonnet 3.5: 59.5%, Physicians: 39.5%). With image integration, all improved, but physicians showed the largest gain (GPT-4o: 84.5%, p<0.001; Claude Sonnet 3.5: 67.3%, p=0.060; Physicians: 78.8%, p<0.001). LLMs changed their explanations in 45-60% of cases when presented with images, demonstrating some level of visual data integration. ConclusionMultimodal LLMs show promise in medical diagnosis, with improved performance when integrating visual evidence. However, this improvement is inconsistent and smaller compared to physicians, indicating a need for enhanced visual data processing in these models.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 95%
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 94%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 94%
Similar papers in this journal
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 93%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
Similar papers in this journal
- GenECG: A synthetic image-based ECG dataset to augment artificial intelligence-enhanced algorithm development 92%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 91%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.