Performance of GPT-4 on the American College of Radiology In-Service Examination
Payne, D. L.; Purohit, K.; Borrero, W. M.; Chung, K.; Hao, M.; Mpoy, M.; Jin, M.; Prasanna, P.; Hill, V.
Show abstract
ObjectivesNo study has evaluated the ability of ChatGPT-4 to answer image-rich diagnostic radiology board exam questions or assessed for model drift in GPT-4s image interpretation abilities. In our study we evaluate GPT-4s performance on the American College of Radiology (ACR) 2022 Diagnostic Radiology In-Training Examination (DXIT). MethodsQuestions were sequentially input into GPT-4 with a standardized prompt. Each answer was recorded and overall accuracy was calculated, as was logic-adjusted accuracy, and accuracy on image-based questions. This experiment was repeated several months later to assess for model drift. ResultsGPT-4 achieved 58.5% overall accuracy, lower than the PGY-3 average (61.9%) but higher than the PGY-2 average (52.8%). Adjusted accuracy was 52.8%. GPT-4 showed significantly higher (p = 0.012) confidence for correct answers (87.1%) compared to incorrect (84.0%). Performance on image-based questions was notably poorer (p < 0.001) at 45.4% compared to text-only questions (80.0%), with adjusted accuracy for image questions of 36.4%. When the questions were repeated, GPT-4 chose a different answer 25.5% of the time and there was a small but insignificant decrease in accuracy. DiscussionGPT-4 performed between PGY-2 and PGY-3 levels on the 2022 DXIT, but significantly poorer on image-based questions, and with large variability in answer choices across time points. This study underscores the potential and risks of using minimally-prompted general AI models in interpreting radiologic images as a diagnostic tool. Implementers of general AI radiology systems should exercise caution given the possibility of spurious yet confident responses.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 93%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 93%
- Classification performance bias between training and test sets in a limited mammography dataset 93%
Similar papers in this journal
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 95%
- Implementation and prospective real-time evaluation of a generalized system for in-clinic deployment and validation of machine learning models in radiology 94%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 94%
Similar papers in this journal
Similar papers in this journal
- Deep learning ensemble for abdominal aortic calcification scoring from lumbar spine X-ray and DXA images 93%
- Two-Step Machine Learning to Diagnose and Predict Involvement of Lungs in COVID-19 and Pneumonia using CT Radiomics 92%
- Fusion of Electronic Health Records and Radiographic Images for a Multimodal Deep Learning Prediction Model of Atypical Femur Fractures 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.