Spot the Difference: Can ChatGPT4-Vision Transform Radiology Artificial Intelligence?
Kelly, B. S. S.; Duignan, S.; Mathur, P.; Dillon, H.; Lee, E.; Yeom, K. W.; Keane, P.; Killeen, R. P.; Lawlor, A.
Show abstract
OpenAIs flagship Large Language Model ChatGPT can now accept image input (GPT4V). "Spot the Difference" and "Medical" have been suggested as emerging applications. The interpretation of medical images is a dynamic process not a static task. Diagnosis and treatment of Multiple Sclerosis is dependent on identification of radiologic change. We aimed to compare the zero-shot performance of GPT4V to a trained U-Net and Vision Transformer (ViT) for the identification of progression of MS on MRI. 170 patients were included. 100 unseen paired images were randomly used for testing. Both U-Net and ViT had 94% accuracy while GPT4V had 85%. GPT4V gave overly cautious non-answers in 6 cases. GPT4V had a precision, recall and F1 score of 0.896, 0.915, 0.905 compared to 1.0, 0.88 and 0.936 for U-Net and 0.94, 0.94, 0.94 for ViT. The impressive performance compared to trained models and a no-code drag and drop interface suggest GPT4V has the potential to disrupt AI radiology research. However misclassified cases, hallucinations and overly cautious non-answers confirm that it is not ready for clinical use. GPT4Vs widespread availability and relatively high error rate highlight the need for caution and education for lay-users, especially those with limited access to expert healthcare. Key pointsO_LIEven without fine tuning and without the need for prior coding experience or additional hardware, GPT4V can perform a zero-shot radiologic change detection task with reasonable accuracy. C_LIO_LIWe find GPT4V does not match the performance of established state of the art computer vision models. GPT4Vs performance metrics are more similar to the vision transformers than the convolutional neural networks, giving some possible insight into its underlying architecture. C_LIO_LIThis is an exploratory experimental study and GPT4V is not intended for use as a medical device. C_LI Summary statementGPT4V can identify radiologic progression of Multiple Sclerosis in a simplified experimental setting. However GPT4V is not a medical device and its widespread availability and relatively high error rate highlight the need for caution and education for lay-users, especially those with limited access to expert healthcare.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 93%
- Brain predictors of fatigue in Rheumatoid Arthritis: a machine learning study 92%
- PRCnet: An Efficient Model for Automatic Detection of Brain Tumor in MRI Images 92%
Similar papers in this journal
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 96%
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 93%
- Impact of Non-Contrast Enhanced Imaging Input Sequences on the Generation of Virtual Contrast-Enhanced Breast MRI Scans using Neural Networks 91%
Similar papers in this journal
- Classification of Hyper-scale Multimodal Imaging Datasets 93%
- Implementation and prospective real-time evaluation of a generalized system for in-clinic deployment and validation of machine learning models in radiology 93%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 91%
Similar papers in this journal
- Radius-Optimized Efficient Template Matching for Lesion Detection from Brain Images 93%
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 93%
- Does contrast-enhancement improve visualisation of lenticulostriate arteries in cerebral small vessel disease using time-of-flight magnetic resonance angiography at 7 Tesla? 92%
Similar papers in this journal
- “This is a quiz” Premise Input: A Key to Unlocking Higher Diagnostic Accuracy in Large Language Models 95%
- Benchmarking Deep Learning-based Image Retrieval of Oral Tumor Histology 89%
- Effects of contrast-medium and vertebral measurement level on computed tomography-based body composition parameters of skeletal muscle and adipose tissue 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.