Robustness Gap of Large Language Models in Nephrology
Soejima, A.; Kitano, F.; Ichikawa, D.; Shibagaki, Y.; Noda, R.
Show abstract
Background: Whether benchmark performance reflects robust clinical reasoning rather than surface-level pattern recognition remains uncertain. We evaluated the robustness of state-of-the-art large language models (LLMs) on nephrology board renewal questions using "None of the other answers" (NOTA) substitution. Methods: From 210 Japanese Society of Nephrology board renewal questions (2014-2023), two nephrologists independently reviewed all items. Questions in which NOTA became the sole correct answer after replacement were included, yielding 145 validated questions. GPT-5, GPT-4o, Gemini 2.5 Pro, and Gemini 2.0 Flash were evaluated via application programming interfaces under default settings. The primary endpoint was accuracy, and paired differences were assessed using the exact two-sided McNemar test. Results: Accuracy was significantly lower after NOTA substitution for all models: GPT-4o, 66.21% to 19.31% (drop, 46.90 percentage points [pp]); GPT-5, 87.59% to 73.10% (14.48 pp); Gemini 2.0 Flash, 58.62% to 31.03% (27.59 pp); and Gemini 2.5 Pro, 86.90% to 55.86% (31.03 pp); all P < .001. GPT-5 showed the smallest decline and the highest accuracy in both versions. Conclusions: All evaluated LLMs showed a significant robustness gap after NOTA replacement. Newer models may be more robust, but multiple-choice accuracy remains an incomplete measure of clinical reasoning robustness.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of a Web-based Prediction Model for Acute Kidney Injury after surgery 93%
- Artificial Intelligence for COVID-19 Risk Classification in Kidney Disease: Can Technology Unmask an Unseen Disease? 93%
- Estimating and predicting kidney function decline in the general population 90%
Similar papers in this journal
- Patient Perceptions by Race of Educational Animations About Living Kidney Donation Made for a Diverse Population 91%
- Evaluating the kidney disease progression using a comprehensive patient profiling algorithm: A hybrid clustering approach 91%
- The Chronic Kidney Disease and Acute Kidney Injury Involvement in COVID-19 Pandemic: A Systematic Review and Meta-analysis 90%
Similar papers in this journal
- Preprint server use in kidney disease research: a rapid review 93%
- The mediating role of trust in physicians on the association between multidimensional health literacy and medication adherence in hemodialysis: A cross-sectional study 93%
- Out-of-sequence placement of deceased donor kidneys is exacerbating inequities in the United States 88%
Similar papers in this journal
- Shared Decision-Making in Renal Replacement Therapy Selection: Patient Perceptions, Preferences, and Influencing Factors in a Nationwide Cross-Sectional Study in Japan 93%
- Development and Validation of a Convolutional Neural Network Model for ICU Acute Kidney Injury Prediction 92%
- A new approach to recognize term and preterm infants with impaired kidney function (IKF) during the first week of life. 90%
Similar papers in this journal
- Ramadan and Kidney disease (RaK) risk assessment tool. Potential Risk Calculator for Evaluating the Risk of Ramadan Fasting In Chronic Kidney Disease patients 92%
- COVID-19 in patients undergoing renal replacement therapy in Scotland: findings and experience from the Scottish Renal Registry 91%
- Superior anticoagulation strategies for renal replacement therapy in critically ill patients with COVID-19: a cohort study 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.