A Reproducible Clinical Decision-Support Suite on MIMIC-IV
Kesan, J. S.
Show abstract
Most published clinical-AI results are single models on a single dataset, difficult to reproduce, and rarely validated outside their training hospital. We built a broad, methodologically rigorous, reproducible clinical decision-support (CDS) suite spanning four families - intensive-care deterioration and outcomes, emergency- department triage, electrocardiographic interpretation, and clini- cal natural-language processing - comprising 26 models. Tabular models are gradient-boosted trees over point-in-time, leakage- safe first-24-hour features; deep models include one-dimensional convolutional networks on raw 12-lead ECG, fine-tuned clinical transformers, and an instruction-tuned large language model for discharge-summary drafting. Every model uses patient-level data splits, probability calibration, a shuffled-label leakage gate, and SHAP explanations, and is characterised by its full confusion matrix with sensitivity, specificity and predictive values. Dis- crimination matched or approached published benchmarks: ICU mortality AUROC 0.884, acute kidney injury 0.830, prolonged stay 0.813; emergency-department-to-ICU 0.875; cardiologist- labelled ECG diagnosis 0.909; full-note diagnostic coding 0.892. Raw-signal ECG deep learning improved myocardial-infarction detection by +0.142 AUROC over interval features. The MIMIC- trained mortality model generalised to a different multi-centre US cohort (199,133 stays) with only a 0.044 AUROC drop. We describe how each model family is incorporated into the latest version of the zMed Critical Care application and its CDS tools
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 95%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 94%
- Developing Machine Learning Models for Predicting Intensive Care Unit Resource Use During the COVID-19 Pandemic 94%
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 94%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 93%
- Multiple Instance Learning Framework can Facilitate Explainability in Murmur Detection 92%
Similar papers in this journal
- Rett syndrome severity estimation with the BioStamp nPoint using interactions between heart rate variability and body movement 92%
- A comparison of machine learning models versus clinical evaluation for mortality prediction in patients with sepsis 92%
- Compressive Big Data Analytics: An Ensemble Meta-Algorithm for High-dimensional Multisource Datasets 91%
Similar papers in this journal
- Real-world evaluation of AI-driven COVID-19 triage for emergency admissions: External validation & operational assessment of lab-free and high-throughput screening solutions 91%
- Remote Covid Assessment in Primary Care (RECAP) risk prediction tool: derivation and real-world validation studies 90%
- CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.