Benchmarking clinical risk prediction algorithms with ensemble machine learning: An illustration of the superlearner algorithm for the non-invasive diagnosis of liver fibrosis in non-alcoholic fatty liver disease
Charu, V.; Liang, J. W.; Mannalithara, A.; Kwong, A.; Tian, L.; Kim, W. R.
Show abstract
Background and AimsEnsemble machine learning (ML) methods can combine many individual models into a single super model using an optimal weighted combination. Here we demonstrate how an underutilized ensemble model, the superlearner, can be used as a benchmark for model performance in clinical risk prediction. We illustrate this by implementing a superlearner to predict liver fibrosis in patients with non-alcoholic fatty liver disease (NAFLD). MethodsWe trained a superlearner based on 23 demographic and clinical variables, with the goal of predicting stage 2 or higher liver fibrosis. The superlearner was trained on data from the Non-alcoholic steatohepatitis - clinical research network observational study (NASH-CRN, n=648), and validated using data from participants in a randomized trial for NASH ( FLINT trial, n=270) and data from examinees with NAFLD who participated in the National Health and Nutrition Examination Survey (NHANES, n=1244). We compared the performance of the superlearner with existing models, including FIB-4, NFS, Forns, APRI, BARD and SAFE. ResultsIn the FLINT and NHANES validation sets, the superlearner (derived from 12 base models) discriminates patients with significant fibrosis from those without well, with AUCs of 0.79 (95% CI: 0.73-0.84) and 0.74 (95% CI: 0.68-0.79). Among the existing scores considered, the SAFE score performed similarly to the superlearner, and the superlearner and SAFE scores outperformed FIB-4, APRI, Forns, and BARD scores in the validation datasets. A superlearner model derived from 12 base models performed as well as one derived from 90 base models. ConclusionsThe superlearner, thought of as the "best-in-class" ML prediction, performed better than most existing models commonly used in practice in detecting fibrotic NASH. The superlearner can be used to benchmark the performance of conventional clinical risk prediction models.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Robust latent-variable interpretation of in vivo regression models by nested resampling 92%
- Multiple instance learning with pathology foundation models effectively predicts kidney disease diagnosis and clinical classification 91%
- DeepMicro: deep representation learning for disease prediction based on microbiome data 91%
Similar papers in this journal
- A user-friendly tool for cloud-based whole slide image segmentation, with examples from renal histopathology 92%
- Deep Proteome Profiling of Metabolic Dysfunction-Associated Steatotic Liver Disease 90%
- Serum Autotaxin is a Prognostic Indicator of Liver-related Events in Patients with Non-alcoholic Fatty Liver Disease 90%
Similar papers in this journal
- RiskPath : Explainable deep learning for multistep biomedical prediction in longitudinal data 91%
- Federated Learning for multi-omics: a performance evaluation in Parkinson's disease 91%
- scTenifoldNet: a machine learning workflow for constructing and comparing transcriptome-wide gene regulatory networks from single-cell data 89%
Similar papers in this journal
- 3D spatially-resolved geometrical and functional models of human liver tissue reveal new aspects of NAFLD progression 92%
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 90%
- Genome-wide polygenic score with APOL1 risk genotypes predicts chronic kidney disease across major continental ancestries 90%
Similar papers in this journal
- Comparative Effectiveness of Medical Concept Embedding for Feature Engineering in Phenotyping 89%
- Framework for Identifying Drug Repurposing Candidates from Observational Healthcare Data 89%
- Clinical interpretation of machine learning models for prediction of diabetic complications using electronic health records 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.