Performance of o1 pro and GPT-4 in self-assessment questions for nephrology board renewal
Noda, R.; Yuasa, C.; Kitano, F.; Ichikawa, D.; Shibagaki, Y.
Show abstract
BackgroundLarge language models (LLMs) are increasingly evaluated in medical education and clinical decision support, but their performance in highly specialized fields, such as nephrology, is not well established. We compared two advanced LLMs, GPT-4 and the newly released o1 pro, on comprehensive nephrology board renewal examinations. MethodsWe administered 209 Japanese Self-Assessment Questions for Nephrology Board Renewal from 2014-2023 to o1 pro and GPT-4 using ChatGPT pro. Each question, including images, was presented in separate chat sessions to prevent contextual carryover. Questions were classified by taxonomy (recall/interpretation/problem-solving), question type (general/clinical), image inclusion, and nephrology subspecialty. We calculated the proportion of correct answers and compared performances using chi-square or Fishers exact tests. ResultsOverall, o1 pro scored 81.3% (170/209), significantly higher than GPT-4s 51.2% (107/209; p<0.001). o1 pro exceeded the 60% passing criterion every year, while GPT-4 achieved this in only two out of the ten years. Across taxonomy levels, question types, and the presence of images, o1 pro consistently outperformed GPT-4 (p<0.05 for multiple comparisons). Performance differences were also significant in several nephrology subspecialties, such as chronic kidney disease, confirming o1 pros broad superiority. Conclusiono1 pro substantially outperformed GPT-4 in a comprehensive nephrology board renewal examination, demonstrating advanced reasoning and integration of specialized knowledge. These findings highlight the potential of next-generation LLMs as valuable tools in specialty medical education and possibly clinical support in nephrology, warranting further and careful validation.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Structural modeling for Oxford histological classifications of immunoglobulin A nephropathy 93%
- Evaluating the kidney disease progression using a comprehensive patient profiling algorithm: A hybrid clustering approach 92%
- Patient Perceptions by Race of Educational Animations About Living Kidney Donation Made for a Diverse Population 92%
Similar papers in this journal
- Development and Validation of a Web-based Prediction Model for Acute Kidney Injury after surgery 94%
- Artificial Intelligence for COVID-19 Risk Classification in Kidney Disease: Can Technology Unmask an Unseen Disease? 93%
- Correlating Deep Learning-Based Automated Reference Kidney Histomorphometry with Patient Demographics and Creatinine 91%
Similar papers in this journal
- The mediating role of trust in physicians on the association between multidimensional health literacy and medication adherence in hemodialysis: A cross-sectional study 93%
- Preprint server use in kidney disease research: a rapid review 92%
- Refining the Composition and Significance of Human Renal Intratubular Casts Using Spatial Protein Imaging 89%
Similar papers in this journal
- Ramadan and Kidney disease (RaK) risk assessment tool. Potential Risk Calculator for Evaluating the Risk of Ramadan Fasting In Chronic Kidney Disease patients 93%
- Health-related hope and reduced distress associated with fluid and dietary restrictions in advanced chronic kidney disease and dialysis: a cohort study 91%
- The Effect of Intradialytic Exercise on Dialysis Patient Survival: A Randomized Controlled Trial 91%
Similar papers in this journal
- Performance of Advanced Large Language Models (GPT-4o, GPT-4, Gemini 1.5 Pro, Claude 3 Opus) on Japanese Medical Licensing Examination: A Comparative Study 94%
- Machine Learning Directed Interventions Associate with Decreased Hospitalization Rates in Hemodialysis Patients 92%
- Predicting mortality in SARS-COV-2 (COVID-19) positive patients in the inpatient setting using a Novel Deep Neural Network 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.