Back

Assessing the Impact of Imputation on the Interpretations of Prediction Models: A Case Study on Mortality Prediction for Patients with Acute Myocardial Infarction

Payrovnaziri, S. N.; Xing, A.; Shaeke, S.; Liu, X.; Bian, J.; He, Z.

2020-09-16 health informatics
10.1101/2020.06.06.20124347 medRxiv
Show abstract

Acute myocardial infarction poses significant health risks and financial burden on healthcare and families. Prediction of mortality risk among AMI patients using rich electronic health record (EHR) data can potentially save lives and healthcare costs. Nevertheless, EHR-based prediction models usually use a missing data imputation method without considering its impact on the performance and interpretability of the model, hampering its real-world applicability in the healthcare setting. This study examines the impact of different methods for imputing missing values in EHR data on both the performance and the interpretations of predictive models. Our results showed that a small standard deviation in root mean squared error across different runs of an imputation method does not necessarily imply a small standard deviation in the prediction models performance and interpretation. We also showed that the level of missingness and the imputation method used can have a significant impact on the interpretation of the models.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.