Back

Performance of reasoning large language models on nephrology multiple-choice questions

Kitano, F.; Masaki, M.; Ichikawa, D.; Shibagaki, Y.; Noda, R.

2025-12-08 nephrology
10.64898/2025.12.03.25341427 medRxiv
Show abstract

AimPerformance of large language models in medicine is improving, yet it remains unclear how the advantage of reasoning models depends on task characteristics in nephrology. MethodsWe evaluated four large language models in two families--OpenAI (GPT-5 reasoning, GPT-4o baseline) and Google (Gemini 2.5 Pro reasoning, Gemini 2.0 Flash baseline)--on 209 self-assessment questions for nephrology board renewal published by the Japanese society of nephrology. Questions were categorized by question type (general vs clinical), taxonomy (recall, interpretation, problem-solving), and image inclusion (non-image vs image). Models were assessed via application programming interface with default parameters; images were provided as PNG files. Accuracy used Wilson 95% confidence intervals (CIs); paired comparisons used McNemars exact test. Primary analyses used logistic generalized linear mixed models with fixed effects, random intercepts, and prespecified interactions. ResultsOverall accuracy was 87.6% (183/209, 95% CI 82.4-91.4) for GPT-5 and 83.7% (175/209, 95% CI 78.1-88.1) for Gemini 2.5 Pro vs. 69.9% (146/209, 95% CI 63.3-75.7) for GPT-4o and 62.7% (131/209, 95% CI 55.9-69.0) for Gemini 2.0 Flash. Paired analyses favored reasoning models, odds ratios of 6.29 for OpenAI and 7.29 for Google (both P<0.001). Adjusted odds ratios for reasoning vs. baseline were 5.00 for OpenAI and 7.28 for Google (both P<0.001). Interactions showed stronger effects in clinical questions for OpenAI and taxonomy-dependent effects for Google; no significant modification by image inclusion. ConclusionReasoning models outperform baseline models with context-dependent advantages in nephrology, although their benefits vary by task and further validation is essential before routine use.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.