Performance of reasoning large language models on nephrology multiple-choice questions
Kitano, F.; Masaki, M.; Ichikawa, D.; Shibagaki, Y.; Noda, R.
Show abstract
AimPerformance of large language models in medicine is improving, yet it remains unclear how the advantage of reasoning models depends on task characteristics in nephrology. MethodsWe evaluated four large language models in two families--OpenAI (GPT-5 reasoning, GPT-4o baseline) and Google (Gemini 2.5 Pro reasoning, Gemini 2.0 Flash baseline)--on 209 self-assessment questions for nephrology board renewal published by the Japanese society of nephrology. Questions were categorized by question type (general vs clinical), taxonomy (recall, interpretation, problem-solving), and image inclusion (non-image vs image). Models were assessed via application programming interface with default parameters; images were provided as PNG files. Accuracy used Wilson 95% confidence intervals (CIs); paired comparisons used McNemars exact test. Primary analyses used logistic generalized linear mixed models with fixed effects, random intercepts, and prespecified interactions. ResultsOverall accuracy was 87.6% (183/209, 95% CI 82.4-91.4) for GPT-5 and 83.7% (175/209, 95% CI 78.1-88.1) for Gemini 2.5 Pro vs. 69.9% (146/209, 95% CI 63.3-75.7) for GPT-4o and 62.7% (131/209, 95% CI 55.9-69.0) for Gemini 2.0 Flash. Paired analyses favored reasoning models, odds ratios of 6.29 for OpenAI and 7.29 for Google (both P<0.001). Adjusted odds ratios for reasoning vs. baseline were 5.00 for OpenAI and 7.28 for Google (both P<0.001). Interactions showed stronger effects in clinical questions for OpenAI and taxonomy-dependent effects for Google; no significant modification by image inclusion. ConclusionReasoning models outperform baseline models with context-dependent advantages in nephrology, although their benefits vary by task and further validation is essential before routine use.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Performance of Advanced Large Language Models (GPT-4o, GPT-4, Gemini 1.5 Pro, Claude 3 Opus) on Japanese Medical Licensing Examination: A Comparative Study 95%
- Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications 91%
- Image and structured data analysis for prognostication of health outcomes in patients presenting to the Emergency Department during the COVID-19 pandemic 90%
Similar papers in this journal
- Evaluation of the performance of GPT-3.5 and GPT-4 on the Medical Final Examination 94%
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 92%
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 92%
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 95%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 93%
Similar papers in this journal
- The mediating role of trust in physicians on the association between multidimensional health literacy and medication adherence in hemodialysis: A cross-sectional study 92%
- Preprint server use in kidney disease research: a rapid review 91%
- Out-of-sequence placement of deceased donor kidneys is exacerbating inequities in the United States 86%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.