A Comparative Study: Diagnostic Performance of ChatGPT 3.5, Google Bard, Microsoft Bing, and Radiologists in Thoracic Radiology Cases
Gunes, Y. C.; Cesur, T.
Show abstract
PurposeTo investigate and compare the diagnostic performance of ChatGPT 3.5, Google Bard, Microsoft Bing, and two board-certified radiologists in thoracic radiology cases published by The Society of Thoracic Radiology. Materials and MethodsWe collected 124 "Case of the Month" from the Society of Thoracic Radiology website between March 2012 and December 2023. Medical history and imaging findings were input into ChatGPT 3.5, Google Bard, and Microsoft Bing for diagnosis and differential diagnosis. Two board-certified radiologists provided their diagnoses. Cases were categorized anatomically (parenchyma, airways, mediastinum-pleura-chest wall, and vascular) and further classified as specific or non-specific for radiological diagnosis. Diagnostic accuracy and differential diagnosis scores were analyzed using chi-square, Kruskal-Wallis and Mann-Whitney U tests. ResultsAmong 124 cases, ChatGPT demonstrated the highest diagnostic accuracy (53.2%), outperforming radiologists (52.4% and 41.1%), Bard (33.1%), and Bing (29.8%). Specific cases revealed varying diagnostic accuracies, with Radiologist I achieving (65.6%), surpassing ChatGPT (63.5%), Radiologist II (52.0%), Bard (39.5%), and Bing (35.4%). ChatGPT 3.5 and Bing had higher differential scores in specific cases (P<0.05), whereas Bard did not (P=0.114). All three had a higher diagnostic accuracy in specific cases (P<0.05). No differences were found in the diagnostic accuracy or differential diagnosis scores of the four anatomical location (P>0.05). ConclusionChatGPT 3.5 demonstrated higher diagnostic accuracy than Bing, Bard and radiologists in text-based thoracic radiology cases. Large language models hold great promise in this field under proper medical supervision.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 97%
- Observer agreement and clinical significance of chest CT reporting in patients suspected of COVID-19 94%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 93%
Similar papers in this journal
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 94%
- Navigated ultrasound bronchoscopy with integrated positron emission tomography - A human feasibility study 94%
- Quantitative analysis of chest computed tomography of COVID-19 pneumonia using a software widely used in Japan 94%
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 97%
- Toward Understanding COVID-19 Pneumonia: A Deep-learning-based Approach for Severity Analysis and Monitoring the Disease 93%
- High-Dimensional Multinomial Multiclass Severity Scoring of COVID-19 Pneumonia Using CT Radiomics Features and Machine Learning Algorithms 93%
Similar papers in this journal
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 95%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 94%
- A Comparison of CXR-CAD Software to Radiologists in Identifying COVID-19 in Individuals Evaluated for Sars CoV 2 Infection in Malawi and Zambia 93%
Similar papers in this journal
- Auto-detection of motion artifacts on CT pulmonary angiograms with a physician-trained AI algorithm 97%
- Volumetric lung cancer screening reduces unnecessary low-dose computed tomography scans: results from a single-centre prospective trial on 4,119 subjects 93%
- A Machine Learning Ensemble Based on Radiomics to Predict BI-RADS Category and Reduce the Biopsy Rate of Ultrasound-Detected Suspicious Breast Masses 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.