Back

Comparative Evaluation of Machine Learning and Deep Learning Models for Early Prediction of Severe Acute Pancreatitis: A Multi-Model Study Using the 2012 Revised Atlanta Classification

stern, N.

2026-06-23 gastroenterology
10.64898/2026.06.20.26356146 medRxiv
Show abstract

**Background:** Acute pancreatitis (AP) is a common gastrointestinal emergency with a subset of patients progressing to severe acute pancreatitis (SAP), which carries substantial morbidity and mortality. Current clinical severity scores such as BISAP, APACHE II, Ranson, and the Modified CT Severity Index require upon 48 hours of observation before reliable assessment is possible, limiting early triage. Machine learning (ML) approaches using routine admission laboratory values may enable earlier, more accurate prediction. **Methods:** We evaluated 11 models spanning three architectural families classical ML (Logistic Regression, Random Forest, Gradient Boosting), feedforward deep learning (MLP, Residual MLP, Attention MLP), and recurrent deep learning (LSTM, Stacked LSTM, Bidirectional LSTM, LSTM+Attention, CNN-LSTM) on a Chinese AP cohort of 722 patients (585 severe, 137 mild) labelled according to the 2012 Revised Atlanta Classification. Performance was assessed via 5-fold stratified cross-validation using AUC-ROC, F1 score, sensitivity, specificity, and PPV, with decision thresholds optimised for maximal F1. **Results:** Random Forest achieved the highest AUC of 0.877 (F1=0.917, sensitivity=96.8%, PPV=87.1%), followed closely by Gradient Boosting (AUC=0.874, F1=0.918). Classical ML models consistently outperformed deep learning counterparts. CNN-LSTM was the best recurrent model (AUC=0.777) but remained inferior to all classical approaches. LSTM-family models produced AUC values of 0.684-0.777, reflecting the cross-sectional tabular nature of the data. **Conclusions:** Random Forest provides robust, high-sensitivity early prediction of SAP severity using routine admission data. External prospective validation is required before clinical deployment. **Keywords:** acute pancreatitis; severity prediction; machine learning; random forest; deep learning; LSTM; Revised Atlanta Classification; early triage

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Diagnostics
50 papers in training set
Top 0.1%
17.1%
2
PLOS ONE
5266 papers in training set
Top 11%
15.7%
3
BMJ Open
601 papers in training set
Top 2%
8.2%
4
Scientific Reports
3612 papers in training set
Top 12%
6.5%
5
American Journal of Gastroenterology
17 papers in training set
Top 0.1%
4.5%
50% of probability mass above
6
Journal of Clinical Medicine
97 papers in training set
Top 1%
3.4%
7
Gut
40 papers in training set
Top 0.3%
3.4%
8
Cureus
68 papers in training set
Top 1%
3.4%
9
PeerJ
308 papers in training set
Top 3%
2.8%
10
Frontiers in Medicine
120 papers in training set
Top 2%
2.1%
11
Heliyon
152 papers in training set
Top 3%
1.8%
12
International Journal of Medical Informatics
26 papers in training set
Top 0.8%
1.6%
13
Contemporary Clinical Trials Communications
11 papers in training set
Top 0.2%
1.5%
14
Nature Communications
5641 papers in training set
Top 48%
1.4%
15
Biomedicines
67 papers in training set
Top 1%
1.4%
16
Annals of Translational Medicine
18 papers in training set
Top 0.3%
1.4%
17
Informatics in Medicine Unlocked
22 papers in training set
Top 0.7%
1.4%
18
Kidney360
22 papers in training set
Top 0.4%
1.4%
19
Journal of Clinical Pathology
15 papers in training set
Top 0.3%
1.4%
20
Medicine
31 papers in training set
Top 1%
1.4%
21
Physiological Measurement
14 papers in training set
Top 0.3%
1.2%
22
Kidney International Reports
15 papers in training set
Top 0.2%
1.2%
23
Biomolecules
100 papers in training set
Top 2%
0.9%
24
Epidemiology and Infection
89 papers in training set
Top 3%
0.6%
25
Frontiers in Pharmacology
111 papers in training set
Top 3%
0.6%
26
Bioengineering & Translational Medicine
21 papers in training set
Top 0.8%
0.5%
27
BMC Cancer
67 papers in training set
Top 3%
0.5%
28
Frontiers in Psychology
56 papers in training set
Top 2%
0.5%
29
PLOS Medicine
110 papers in training set
Top 4%
0.5%
30
F1000Research
88 papers in training set
Top 5%
0.5%