Accuracy of Artificial Intelligence-Based Models versus Traditional Scoring Systems (APACHE, SOFA, SAPS) for Predicting Mortality in ICU Patients: A Systematic Review and Meta-Analysis
Pradhan, V.; SHEKHAR, H.; Munda, P. K.; Tiwari, A. K.; Jha, S.; Rai, P.
Show abstract
IntroductionReliable estimation of mortality among critically ill patients is crucial for guiding clinical decisions and optimizing ICU performance. Traditional scoring systems such as APACHE, SOFA, and SAPS are commonly applied, though their predictive capacity is constrained by their reliance on static structures and linear modeling assumptions. Artificial intelligence-based models provide flexible, data-oriented prediction strategies, yet their comparative accuracy remains unclear. This study systematically reviewed and meta-analyzed the performance of Artificial intelligence-based models versus conventional ICU scores for predicting in-hospital mortality in adults admitted to ICU. Materials and MethodsLiterature searches were performed in PubMed, Embase, Web of Science, Scopus and the Cochrane Library from January 2015 to August 2025 for studies comparing AI models with traditional scoring systems. Studies were included if they provided diagnostic performance indicators including AUC, sensitivity, or specificity. Risk of bias was assessed using PROBAST, and pooled statistical estimates were derived through bivariate random-effects modeling with Fisher s Z- transformation. Subgroup analyses examined AI modality, ICU type, and geographic region. ResultsEleven studies involving over one million ICU admissions met inclusion criteria. Two studies (Huang 2023; Lim 2024) provided complete 2 by 2 data for meta-analysis. Pooled sensitivity and specificity for AI models were 0.875 (95% CI: 0.840 - 0.904) and 0.857 (95% CI: 0.845 - 0.868), respectively. AI models achieved higher AUCs (0.82 - 0.90) than APACHE II (0.70 - 0.78), SOFA (0.68 - 0.75), and SAPS II (0.70 - 0.79). Deep learning and ensemble methods performed best across ICU settings and regions. ConclusionAI-based models outperform conventional scoring systems in predicting ICU mortality. Their integration into critical care could enhance early risk stratification and precision prognostication. HighlightsThis meta-analysis highlights that artificial intelligence-based predictive models demonstrated superior predictive performance than conventional ICU scoring systems (APACHE, SOFA, and SAPS) in predicting in-hospital mortality, with higher pooled sensitivity, specificity, and overall discriminative accuracy, particularly for deep learning and ensemble approaches across diverse ICU settings.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The COVID-19 Critical Care Consortium observational study: Design and rationale of a prospective, international, multicenter, observational study 94%
- Development and validation of automated computer aided-risk score for predicting the risk of in-hospital mortality using first electronically recorded blood test results and vital signs for COVID-19 hospital admissions: a retrospective development and validation study 94%
- Identification of Acute Respiratory Distress Syndrome subphenotypes denovo using routine clinical data: a retrospective analysis of ARDS clinical trials 93%
Similar papers in this journal
- COVID-19 ICU and mechanical ventilation patient characteristics and outcomes - A systematic review and meta-analysis 94%
- Can we predict the severe course of COVID-19 – a systematic review and meta-analysis of indicators of clinical outcome? 94%
- Derivation and validation of a triage tool for acutely ill adults with suspected COVID-19: The PRIEST observational cohort study 94%
Similar papers in this journal
- ABCDEF Bundle Implementation: The influence of access to bundle-enhancing supplies and equipment 93%
- Acute respiratory distress syndrome and shunt detection with bubble studies: a systematic review and meta-analysis 92%
- Divergence Between Net Fluid and Weight-Based Evaluation in Calculating Cumulative Fluid Balance 91%
Similar papers in this journal
- Imputation of PaO2 from SpO2 values from the MIMIC-III Critical Care Database Using Machine-Learning Based Algorithms 93%
- AKI Risk Score (AKI-RiSc): Developing an Interpretable Clinical Score for Early Identification of Acute Kidney Injury for Patients Presenting to the Emergency Department 92%
- Machine learning vs. traditional regression analysis for fluid overload prediction in the ICU 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.