From Clinical Judgement to Large Language Models: Benchmarking Predictive Approaches for Unplanned Hospital Admissions
Neves, B.; Silva, M. J.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSBackgroundC_ST_ABSWhile machine learning (ML) models show strong performance for predicting unplanned hospital visits, their clinical utility relative to physician judgment remains unclear. Large language models (LLMs) offer a promising middle ground, potentially combining algorithmic accuracy with human-interpretable reasoning. ObjectiveTo directly compare the predictive performance of physicians, structured ML models, and LLMs for forecasting 30-day emergency department (ED) visits and unplanned hospital admissions under equivalent data conditions. MethodsWe selected 404 cases from structured EHR data and converted them into synthetic clinical vignettes using GPT-5. Thirty-five physicians evaluated these vignettes, while CLMBR-T (a machine learning model trained on structured EHR data) was applied to the original data. Eight LLMs evaluated the same vignettes. We compared discriminative performance (AUROC, AUPRC), calibration (Brier score, Expected Calibration Error), and confidence-performance relationships across all methods. ResultsCLMBR-T achieved the highest discriminative performance (AUROC 0.79, 95% CI: 0.75-0.83; AUPRC 0.78, 95% CI: 0.72-0.83), followed by large LLMs (DeepSeek V3, Claude 4.1 Opus, GPT-5; AUROC 0.74). Pooled physicians performed lowest (AUROC 0.65, 95% CI: 0.59-0.70; AUPRC 0.61, 95% CI: 0.54-0.68). However, LLMs showed stronger alignment with physician reasoning (correlation r=0.51-0.65) compared to CLMBR-T (r=0.37). CLMBR-T demonstrated superior confidence calibration with significant confidence-performance correlation (r=0.21, p<0.001), while physicians showed poor calibration (r=0.07, p=0.17). Individual physician performance varied widely (AUROC 0.55-0.83), with three out of 35 physicians exceeding the ML benchmark. ConclusionsML models trained on structured EHR data outperform both physicians and LLMs in predictive accuracy and confidence calibration, though LLMs achieved competitive zero-shot performance and better approximated human clinical reasoning. These findings suggest hybrid approaches combining high-performance ML screening with interpretable LLM explanations may optimize both accuracy and clinical adoption. The substantial variability in physician performance highlights limitations of benchmarking against "average" clinical judgment.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 95%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 94%
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 93%
Similar papers in this journal
- Use of unstructured text in prognostic clinical prediction models: a systematic review 95%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 95%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 95%
Similar papers in this journal
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 94%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 93%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 94%
- Implicit bias in Critical Care Data: Factors affecting sampling frequencies and missingness patterns of clinical and biological variables in ICU Patients 93%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 93%
Similar papers in this journal
- Modeling physician variability to prioritize relevant medical record information 95%
- A deep learning model for clinical outcome prediction using longitudinal inpatient electronic health records 93%
- Evaluation of Patient-Level Retrieval from Electronic Health Record Data for a Cohort Discovery Task 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.