Comparative Performance of Claude and GPT Models in Basic Radiological Imaging Tasks
Nguyen, C.; Carrion, D.; Badawy, M.
Show abstract
BackgroundPublicly available artificial intelligence (AI) Visual Language Models (VLMs) are constantly improving. The advent of vision capabilities on these models could enhance workflows in radiology. Evaluating their performance in radiological image interpretation is vital to their potential integration into practice. AimThis study aims to evaluate the proficiency and consistency of the publicly available VLMs, Claude and GPT, across multiple iterations in basic image interpretation tasks. MethodSubsets from publicly available datasets, ROCOv2 and MURAv1.1, were used to evaluate 6 VLMs. A system prompt and image were inputted into each model thrice. The outputs were compared to the dataset captions to evaluate each modelss accuracy in recognising the modality, anatomy, and detecting fractures on radiographs. The consistency of the output across iterations was also analysed. ResultsEvaluation of the ROCOv2 dataset showed high accuracy in modality recognition, with some models achieving 100%. Anatomical recognition ranged between 61% and 85% accuracy across all models tested. On the MURAv1.1 dataset, Claude-3.5-Sonnet had the highest anatomical recognition with 57% accuracy, while GPT-4o had the best fracture detection with 62% accuracy. Claude-3.5-Sonnet was the most consistent model, with 83% and 92% consistency in anatomy and fracture detection, respectively. ConclusionGiven Claude and GPTs current accuracy and reliability, integration of these models into clinical settings is not yet feasible. This study highlights the need for ongoing development and establishment of standardised testing techniques to ensure these models achieve reliable performance.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Classification of Hyper-scale Multimodal Imaging Datasets 96%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 95%
- Implementation and prospective real-time evaluation of a generalized system for in-clinic deployment and validation of machine learning models in radiology 94%
Similar papers in this journal
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 95%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 94%
- Investigating the relationship between internal spinal alignment and back shape in patients with scoliosis using PCdare: a comparative, reliability and validation study 94%
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 97%
- Fully automatic segmentation of craniomaxillofacial CT scans for computer-assisted orthognathic surgery planning using the nnU-Net framework 93%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 92%
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 94%
- MyoVision-US: an Artificial Intelligence-Powered Software for Automated Analysis of Skeletal Muscle Ultrasonography 93%
- Effective Deep Learning Approaches for Predicting COVID-19 Outcomes from Chest Computed Tomography Volumes 93%
Similar papers in this journal
- Fully Automated Explainable Abdominal CT Contrast Media Phase Classification Using Organ Segmentation and Machine Learning 95%
- Comparing performance of deep convolutional neural network with orthopaedic surgeons on identification of total hip prosthesis design from plain radiographs 93%
- Phase Recognition in Contrast-Enhanced CT Scans based on Deep Learning and Random Sampling 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.