Reproducibility and Robustness of Localized Mortality Prediction
Nitch-Griffin, E.; Peterson, A.; Skaf, Y.; Brunson, J. C.
Show abstract
BackgroundWhile localized modeling--the use of predictive models to perform the adaptation step in case-based reasoning--has been evaluated in several experimental settings, its reported successes have infrequently been independently and externally validated. ObjectiveWe aimed to extend and validate an experimental study of mortality prediction in a critical care population and to assess the importance of several methodological factors to predictive performance. MethodsWe reproduced the workflow of Lee, Maslove, and Dubin (2015) using an updated database. We evaluated performance as area under the receiver operating characteristic curve and under the precision-recall curve and calibration as weakness of evidence in Hosmer- Lemeshow tests. We compared the effects of several modeling choices, including how relevance is quantified, and how relevance cohorts are retrieved, and the choice of model. We compared ours to previous results and used linear regression to quantify the role of each modeling choice on performance. ResultsOverall performance and its relationship to cohort size validated previous results. These relationships varied by model family as expected, though we observed no advantage of decision trees over a model-free approach and poor performance by random forests. An alternate choice of unlearned similarity measure yielded marginal and inconsistent performance differences. Denominating cohorts by similarity threshold rather than by cardinality yielded marginal but consistent performance losses. A temporal validation exercise enabled by a change of information system before the recent upgrade corroborated performance estimates from cross-validation. In all, cohort denomination mattered more to performance than any other methodological choice. DiscussionThe greater impact of retrieval than of adaptation suggests a weakness with the strategy of localized modeling. Additional research to deconstruct the varieties of this approach and quantify the relative benefits of its components is needed to resolve this question.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Implicit bias in Critical Care Data: Factors affecting sampling frequencies and missingness patterns of clinical and biological variables in ICU Patients 95%
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 94%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 94%
Similar papers in this journal
Similar papers in this journal
- Modeling physician variability to prioritize relevant medical record information 95%
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 94%
- Evaluation of Patient-Level Retrieval from Electronic Health Record Data for a Cohort Discovery Task 92%
Similar papers in this journal
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 93%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 92%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.