Generalisability of AI-based scoring systems in the ICU: a systematic review and meta-analysis
Rockenschaub, P.; Akay, E. M.; Carlisle, B. G.; Hilbert, A.; Meyer-Eschenbach, F.; Näher, A.-F.; Frey, D.; Madai, V. I.
Show abstract
BackgroundMachine learning (ML) is increasingly used to predict clinical deterioration in intensive care unit (ICU) patients through scoring systems. Although promising, such algorithms often overfit their training cohort and perform worse at new hospitals. Thus, external validation is a critical - but frequently overlooked - step to establish the reliability of predicted risk scores to translate them into clinical practice. We systematically reviewed how regularly external validation of ML-based risk scores is performed and how their performance changed in external data. MethodsWe searched MEDLINE, Web of Science, and arXiv for studies using ML to predict deterioration of ICU patients from routine data. We included primary research published in English before April 2022. We summarised how many studies were externally validated, assessing differences over time, by outcome, and by data source. For validated studies, we evaluated the change in area under the receiver operating characteristic (AUROC) attributable to external validation using linear mixed-effects models. ResultsWe included 355 studies, of which 39 (11.0%) were externally validated, increasing to 17.9% by 2022. Validated studies made disproportionate use of open-source data, with two well-known US datasets (MIMIC and eICU) accounting for 79.5% of studies. On average, AUROC was reduced by -0.037 (95% CI -0.064 to -0.017) in external data, with >0.05 reduction in 38.6% of studies. DiscussionExternal validation, although increasing, remains uncommon. Performance was generally lower in external data, questioning the reliability of some recently proposed ML-based scores. Interpretation of the results was challenged by an overreliance on the same few datasets, implicit differences in case mix, and exclusive use of AUROC.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Can we predict the severe course of COVID-19 – a systematic review and meta-analysis of indicators of clinical outcome? 95%
- Population risk factors for severe disease and mortality in COVID-19: A global systematic review and meta-analysis 95%
- Regional performance variation in external validation of four prediction models for severity of COVID-19 at hospital admission: An observational multi-centre cohort study 95%
Similar papers in this journal
- Development and validation of automated computer aided-risk score for predicting the risk of in-hospital mortality using first electronically recorded blood test results and vital signs for COVID-19 hospital admissions: a retrospective development and validation study 96%
- Use of the first National Early Warning Score recorded within 24 hours of admission to estimate the risk of in-hospital mortality in unplanned COVID-19 patients: a retrospective cohort study 94%
- Identification of Acute Respiratory Distress Syndrome subphenotypes denovo using routine clinical data: a retrospective analysis of ARDS clinical trials 94%
Similar papers in this journal
Similar papers in this journal
- Use of unstructured text in prognostic clinical prediction models: a systematic review 96%
- Validation of a Derived International Patient Severity Algorithm to Support COVID-19 Analytics from Electronic Health Record Data 94%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.