Boundary-Specific Failure Modes and Safety Trade-offs of Large Language Models in ChronicKidney Disease Renoprotective Therapy Review:A Stratified Synthetic Benchmark
Yeh, S.-E.; Lin, H.-J.; Lai, W.-W.; Lin, H.
Show abstract
Background.Renoprotective therapies - SGLT2 inhibitors, finerenone, and renin-angiotensin system inhibitors (RASi) - remain underutilisedin chronic kidney disease (CKD). Large language models (LLMs) may detect therapy omissions, but their performance acrossCKD severity strata and at clinical decision boundaries has not been evaluated.Methods.We constructed 100 synthetic CKD vignettes (G3a-G5D; 75 with prespecified omissions, 25 decoys) and queried four LLMsthree times each at temperature 0 (1,200 calls). Omission criteria were adapted from KDIGO 2024, including an investigator-defined gray-zone RASi initiation criterion at eGFR<15. Two nephrologists independently classified a stratified 20-casesubset.Results.For SGLT2 inhibitor and finerenone omissions, all models achieved near-ceiling sensitivity (97-100%). For RASi, performancediverged at the eGFR<15 boundary: Grok 4.1 Fast 85% versus GPT-5.4 55%, Gemini 10%, DeepSeek 10%. Gap-detectioninter-rater agreement was perfect (kappa = 1.000). Clinically incorrect reasoning rates ranged from 0% (GPT-5.4) to 27%(DeepSeek R1); of 52 instances, 31 were factual pharmacology errors and 21 reflected conservative boundary-discordantreasoning. Reproducibility (Jaccard) ranged from 0.74 to 0.93.Conclusions.This boundary-aware synthetic benchmark showed that aggregate sensitivity can conceal clinically important operational-rulediscordance. Rule-based SGLT2 inhibitor and finerenone omissions were detected with near-ceiling sensitivity, whereas aninvestigator-defined gray-zone RASi criterion at eGFR<15 exposed model-specific boundary behaviour. Evaluation of LLM-based CKD decision support should report boundary-specific performance, reproducibility, and clinically incorrect reasoningalongside aggregate metrics.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Preprint server use in kidney disease research: a rapid review 91%
- Comparison of low eGFR prevalence and prediction for mortality using 2009 and 2021 CKD-EPI equations in Mexican adults 90%
- The mediating role of trust in physicians on the association between multidimensional health literacy and medication adherence in hemodialysis: A cross-sectional study 90%
Similar papers in this journal
- ATP-citrate lyase as a therapeutic target in chronic kidney disease: a Mendelian Randomization analysis 92%
- The prevalence of chronic kidney disease in Australian primary care: analysis of a national general practice dataset 91%
- Shared Decision-Making in Renal Replacement Therapy Selection: Patient Perceptions, Preferences, and Influencing Factors in a Nationwide Cross-Sectional Study in Japan 89%
Similar papers in this journal
- Detailed disease progression of 213 patients hospitalized with Covid-19 in the Czech Republic: An exploratory analysis 91%
- Effect of common maintenance drugs on the risk and severity of COVID-19 in elderly patients 91%
- A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program 91%
Similar papers in this journal
Similar papers in this journal
- COVID-19 in patients undergoing renal replacement therapy in Scotland: findings and experience from the Scottish Renal Registry 92%
- External validation of six clinical models for prediction of unknown chronic kidney disease in a German population 91%
- Seasonality of acute kidney injury phenotypes in England: an unsupervised machine learning classification study of electronic health records 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.