Evaluation of Closed and Open Large Language Models in Pediatric Cardiology Board Exam Performance
Nikolovski, N.; Morgan, C. T.; Gritti, M. N.
Show abstract
IntroductionLarge language models (LLMs) have gained traction in medicine, but there is limited research comparing closed- and open-source models in subspecialty contexts. This study evaluated ChatGPT-4.0o and DeepSeek-R1 on a pediatric cardiology board-style examination to quantify their accuracy and discuss clinical and educational utility. MethodsChatGPT-4.0o and DeepSeek-R1 were used to answer 88 text-based multiple-choice questions across 11 pediatric cardiology subtopics from a Pediatric Cardiology Board Review textbook. DeepSeek-R1s processing time per question was measured. Statistical analyses for model comparison were conducted using an unpaired two-tailed t-test, and bivariate correlations were assessed using Pearsons r. ResultsChatGPT-4.0o and DeepSeek-R1 achieved 70% (62/88) and 68% (60/88) accuracy, respectively (p=0.79). Subtopic accuracy was equal in 5 of 11 chapters, with each model outperforming its counterpart in 3 of 11. DeepSeek-R1s processing time negatively correlated with accuracy (r = -0.68, p = 0.02). ConclusionChatGPT-4.0o and DeepSeek-R1 approached the passing threshold on a pediatric cardiology board examination, with comparable accuracy and potential for open-source models to enhance clinical and educational outcomes while supporting sustainable AI development.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 94%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 93%
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 92%
Similar papers in this journal
- The Impact of COVID-19 Pandemic on Cardiology Services 92%
- Machine learning approaches to predict 30-day mortality following percutaneous coronary intervention in an Australian population 92%
- Impact of COVID-19 pandemic on rates of congenital heart disease procedures among children: Prospective cohort analyses of 26,270 procedures in 17,860 children using CVD-COVID-UK consortium record linkage data 91%
Similar papers in this journal
- GenECG: A synthetic image-based ECG dataset to augment artificial intelligence-enhanced algorithm development 93%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 90%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 89%
Similar papers in this journal
- ChatGPT Provides Inconsistent Risk-Stratification of Patients With Atraumatic Chest Pain 94%
- Multimodal Prehabilitation in People Awaiting Acute Inpatient Cardiac Surgery: Study Protocol for a Pilot Feasibility Trial (PreP-ACe) 91%
- Validation of deep learning enabled web based and smartphone optimized application RadAnalyzer to measure vertebral heart size and vertebral left atrial size in dogs 91%
Similar papers in this journal
- Hospitalist perspectives of available tests to monitor volume status in patients with heart failure: a qualitative study 91%
- Referral for Cardiac Amyloidosis in Patients who underwent Transcatheter Aortic Valve Replacement: Result of Quality Outcome Project 90%
- Evaluation of Self-Directed Learning Activities at King Abdulaziz University: A Qualitative Study of Faculty Perceptions 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.