Algorithmic Versus Expert Rankings of Large Language Models in Peritoneal Dialysis Prescription Review: A Trap-Embedded Synthetic Benchmark
Wei, C.-H.; Lin, H.-J.; Lai, W.-W.; Lin, H. M.
Show abstract
Background: Clinical LLM benchmarks rarely test whether algorithmic rankings agree with expert clinical judgment. We developed a trap-embedded peritoneal dialysis (PD) benchmark comparing multiple scoring constructs with blinded nephrologist ratings. Methods: We generated 125 synthetic PD cases containing 13 ISPD-aligned trap types. Five LLMs (Claude Sonnet 4.5, GPT-5.4, Gemini 3.1 Pro, DeepSeek-R1, Grok 4.1 Fast) evaluated each case three times at temperature 0 (1,875 calls). Primary outcome was must-identify TDR_must, analyzed with GEE and case-clustered bootstrap. Secondary analyses included a verbosity-sensitive alarm-burden proxy, WCS, relaxed-match scoring, WCS sensitivity analyses, and a 25-output blinded expert adequacy substudy. Must-identify kappa was 0.89 in Stage 1 and 0.92 in Stage 2. Results: Rankings were discordant. Recall ranked Claude (0.977) and GPT-5.4 (0.955) above the other models (0.86-0.90, p<0.0001). The alarm-burden proxy favored concise models (Grok 0.689; 21.6 vs 2.4 issues/case), while WCS produced a third ordering. In the expert substudy, inter-rater concordance was strong (rho 0.977), but WCS did not show a positive association with expert adequacy (rho -0.17, p=0.41). Conclusion: Clinical LLM rankings in PD prescription review depend strongly on scoring construct. Algorithmic metrics should be reported alongside blinded expert adequacy ratings and should not alone determine deployment.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program 91%
- Evaluating the kidney disease progression using a comprehensive patient profiling algorithm: A hybrid clustering approach 90%
- Efficacy of Tenapanor in Managing Hyperphosphatemia and Constipation in Hemodialysis Patients: A Randomized Controlled Trial 89%
Similar papers in this journal
- Shared Decision-Making in Renal Replacement Therapy Selection: Patient Perceptions, Preferences, and Influencing Factors in a Nationwide Cross-Sectional Study in Japan 90%
- ATP-citrate lyase as a therapeutic target in chronic kidney disease: a Mendelian Randomization analysis 90%
- The prevalence of chronic kidney disease in Australian primary care: analysis of a national general practice dataset 90%
Similar papers in this journal
- Preprint server use in kidney disease research: a rapid review 91%
- The mediating role of trust in physicians on the association between multidimensional health literacy and medication adherence in hemodialysis: A cross-sectional study 89%
- Comparison of low eGFR prevalence and prediction for mortality using 2009 and 2021 CKD-EPI equations in Mexican adults 89%
Similar papers in this journal
- Correlating Deep Learning-Based Automated Reference Kidney Histomorphometry with Patient Demographics and Creatinine 91%
- Artificial Intelligence for COVID-19 Risk Classification in Kidney Disease: Can Technology Unmask an Unseen Disease? 90%
- Estimating and predicting kidney function decline in the general population 90%
Similar papers in this journal
- Clonal hematopoiesis of indeterminate potential contributes to accelerated chronic kidney disease progression 91%
- Prognostic Utility of Total Kidney Volume for Chronic Kidney Disease Risk Prediction: An Observational and Mendelian Randomization Study 90%
- Heterogeneous treatment effects of intensive glycemic control on kidney microvascular outcomes in ACCORD 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.