Comparison of the diagnostic accuracy among GPT-4 based ChatGPT, GPT-4V based ChatGPT, and radiologists in musculoskeletal radiology
Horiuchi, D.; Tatekawa, H.; Oura, T.; Shimono, T.; Walston, S. L.; Takita, H.; Matsushita, S.; Mitsuyama, Y.; Miki, Y.; Ueda, D.
Show abstract
ObjectiveTo compare the diagnostic accuracy of Generative Pre-trained Transformer (GPT)-4 based ChatGPT, GPT-4 with vision (GPT-4V) based ChatGPT, and radiologists in musculoskeletal radiology. Materials and MethodsWe included 106 "Test Yourself" cases from Skeletal Radiology between January 2014 and September 2023. We input the medical history and imaging findings into GPT-4 based ChatGPT and the medical history and images into GPT-4V based ChatGPT, then both generated a diagnosis for each case. Two radiologists (a radiology resident and a board-certified radiologist) independently provided diagnoses for all cases. The diagnostic accuracy rates were determined based on the published ground truth. Chi-square tests were performed to compare the diagnostic accuracy of GPT-4 based ChatGPT, GPT-4V based ChatGPT, and radiologists. ResultsGPT-4 based ChatGPT significantly outperformed GPT-4V based ChatGPT (p < 0.001) with accuracy rates of 43% (46/106) and 8% (9/106), respectively. The radiology resident and the board-certified radiologist achieved accuracy rates of 41% (43/106) and 53% (56/106). The diagnostic accuracy of GPT-4 based ChatGPT was comparable to that of the radiology resident but was lower than that of the board-certified radiologist, although the differences were not significant (p = 0.78 and 0.22, respectively). The diagnostic accuracy of GPT-4V based ChatGPT was significantly lower than those of both radiologists (p < 0.001 and < 0.001, respectively). ConclusionGPT-4 based ChatGPT demonstrated significantly higher diagnostic accuracy than GPT-4V based ChatGPT. While GPT-4 based ChatGPTs diagnostic performance was comparable to radiology residents, it did not reach the performance level of board-certified radiologists in musculoskeletal radiology.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 97%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 92%
- Observer agreement and clinical significance of chest CT reporting in patients suspected of COVID-19 92%
Similar papers in this journal
- Investigating the relationship between internal spinal alignment and back shape in patients with scoliosis using PCdare: a comparative, reliability and validation study 94%
- Classification performance bias between training and test sets in a limited mammography dataset 93%
- Navigated ultrasound bronchoscopy with integrated positron emission tomography - A human feasibility study 93%
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 96%
- Image matters: Development of a novel staining method, the RGB-trichrome, as a reliable tool for the study of the musculoskeletal system 93%
- MyoVision-US: an Artificial Intelligence-Powered Software for Automated Analysis of Skeletal Muscle Ultrasonography 92%
Similar papers in this journal
- “This is a quiz” Premise Input: A Key to Unlocking Higher Diagnostic Accuracy in Large Language Models 95%
- Effects of contrast-medium and vertebral measurement level on computed tomography-based body composition parameters of skeletal muscle and adipose tissue 93%
- Benchmarking Deep Learning-based Image Retrieval of Oral Tumor Histology 92%
Similar papers in this journal
- Fully Automated Explainable Abdominal CT Contrast Media Phase Classification Using Organ Segmentation and Machine Learning 94%
- Comparing performance of deep convolutional neural network with orthopaedic surgeons on identification of total hip prosthesis design from plain radiographs 91%
- Phase Recognition in Contrast-Enhanced CT Scans based on Deep Learning and Random Sampling 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.