Empowering Radiologists with ChatGPT-4o: Comparative Evaluation of Large Language Models and Radiologists in Cardiac Cases
Cesur, T.; Gunes, Y. C.; Camur, E.; Dagli, M.
Show abstract
PurposeThis study evaluated the diagnostic accuracy and differential diagnosis capabilities of 12 Large Language Models (LLMs), one cardiac radiologist, and three general radiologists in cardiac radiology. The impact of ChatGPT-4o assistance on radiologist performance was also investigated. Materials and MethodsWe collected publicly available 80 "Cardiac Case of the Month from the Society of Thoracic Radiology website. LLMs and Radiologist-III were provided with text-based information, whereas other radiologists visually assessed the cases with and without ChatGPT-4o assistance. Diagnostic accuracy and differential diagnosis scores (DDx Score) were analyzed using the chi-square, Kruskal-Wallis, Wilcoxon, McNemar, and Mann-Whitney U tests. ResultsThe unassisted diagnostic accuracy of the cardiac radiologist was 72.5%, General Radiologist-I was 53.8%, and General Radiologist-II was 51.3%. With ChatGPT-4o, the accuracy improved to 78.8%, 70.0%, and 63.8%, respectively. The improvements for General Radiologists-I and II were statistically significant (P[≤]0.006). All radiologists DDx scores improved significantly with ChatGPT-4o assistance (P[≤]0.05). Remarkably, Radiologist-Is GPT-4o-assisted diagnostic accuracy and DDx Score were not significantly different from the Cardiac Radiologists unassisted performance (P>0.05). Among the LLMs, Claude 3.5 Sonnet and Claude 3 Opus had the highest accuracy (81.3%), followed by Claude 3 Sonnet (70.0%). Regarding the DDx Score, Claude 3 Opus outperformed all models and Radiologist-III (P<0.05). The accuracy of the general radiologist-III significantly improved from 48.8% to 63.8% with GPT4o-assistance (P<0.001). ConclusionChatGPT-4o may enhance the diagnostic performance of general radiologists for cardiac imaging, suggesting its potential as a valuable diagnostic support tool. Further research is required to assess its clinical integration.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 96%
- Observer agreement and clinical significance of chest CT reporting in patients suspected of COVID-19 93%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 92%
Similar papers in this journal
- Accuracy of deep learning based computed tomography diagnostic system of COVID-19: a consecutive sampling external validation cohort study 94%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 94%
- Navigated ultrasound bronchoscopy with integrated positron emission tomography - A human feasibility study 93%
Similar papers in this journal
- Evaluation of the second-generation whole-heart motion correction algorithm (SSF2) used to demonstrate the aortic annulus on cardiac CT 95%
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 95%
- Toward Understanding COVID-19 Pneumonia: A Deep-learning-based Approach for Severity Analysis and Monitoring the Disease 94%
Similar papers in this journal
- Auto-detection of motion artifacts on CT pulmonary angiograms with a physician-trained AI algorithm 96%
- Detection, Isolation and Quantification of Myocardial Infarct with Four Different Histological Staining Techniques 93%
- Demarcation line determination for diagnosis of gastric cancer disease range using unsupervised machine learning in magnifying narrow-band imaging 91%
Similar papers in this journal
- “This is a quiz” Premise Input: A Key to Unlocking Higher Diagnostic Accuracy in Large Language Models 95%
- Effects of contrast-medium and vertebral measurement level on computed tomography-based body composition parameters of skeletal muscle and adipose tissue 93%
- Post mortem pathological findings in COVID-19 cases: A Systematic Review 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.