Evaluating AI Reasoning Models in Pediatric Medicine: A Comparative Analysis of o3-mini and o3-mini-high
Mondillo, G.; Masino, M.; Colosimo, S.; Perrotta, A.; Frattolillo, V.
Show abstract
Artificial intelligence (AI) is increasingly playing a crucial role in modern medicine, particularly in clinical decision support. This study compares the performance of two OpenAI reasoning models, o3-mini and o3-mini-high, in answering 900 pediatric clinical questions derived from the MedQA-USMLE dataset. The evaluation focuses on accuracy, response time, and consistency to determine their effectiveness in pediatric diagnostic and therapeutic decision-making. The results indicate that o3-mini-high achieves a higher accuracy (90.55% vs. 88.3%) and faster response times (64.63 seconds vs. 71.63 seconds) compared to o3-mini. The chi-square test confirmed that these differences are statistically significant (X2 = 328.9675, p < 0.00001)). Error analysis revealed that o3-mini-high corrected more errors from o3-mini than vice versa, but both models shared 61 common errors, suggesting intrinsic limitations in training data or model architecture. Additionally, accessibility differences between the models were considered. While DeepSeek-R1, evaluated in a previous study, offers unrestricted free access, OpenAIs o3 models have message limitations, potentially influencing their suitability in resource-constrained environments. Future improvements should aim at reducing shared errors, optimizing o3-minis accuracy while maintaining efficiency, and refining o3-mini-high for enhanced performance. Implementing an ensemble approach that leverages both models strengths could provide a more robust AI-driven clinical decision support system, particularly in time-sensitive pediatric scenarios such as emergency care and neonatal intensive care units.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine Learning for Paediatric Related Decision Support in Emergency Care - A UK and Ireland Network Survey Study 94%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 93%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
Similar papers in this journal
- A Machine Learning-Based Prediction of Hospital Mortality in Mechanically Ventilated ICU Patients 94%
- ChatGPT-Enhanced ROC Analysis (CERA): A Shiny Web Tool for Finding Optimal Cutoff in Biomarker Analysis 93%
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 92%
Similar papers in this journal
- Prediction of COVID-19 Mortality to Support Patient Prognosis and Triage and Limits of Current Open-Source Data 91%
- Caregivers’ Perspective: Satisfaction With Healthcare Services At The Paediatric Specialist Clinic Of The National Referral Centre In Malaysia 91%
- Learning from the resilience of hospitals and their staff to the COVID-19 pandemic: a scoping review 90%
Similar papers in this journal
- Evaluating Explanations from AI Algorithms for Clinical Decision-Making: A Social Science-based Approach 94%
- Deep Sentiment Classification and Topic Discovery on Novel Coronavirus or COVID-19 Online Discussions: NLP Using LSTM Recurrent Neural Network Approach 91%
- Off-body Sleep Analysis for Predicting Adverse Behavior in Individuals with Autism Spectrum Disorder 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.