A framework for testing different imputation methods for tabular datasets
Kossen, T.; Livne, M.; Madai, V. I.; Galinovic, I.; Frey, D.; Fiebach, J. B.
Show abstract
Background and purposeHandling missing values is a prevalent challenge in the analysis of clinical data. The rise of data-driven models demands an efficient use of the available data. Methods to impute missing values are thus crucial. Here, we developed a publicly available framework to test different imputation methods and compared their impact in a typical stroke clinical dataset as a use case.\n\nMethodsA clinical dataset based on the 1000Plus stroke study with 380 completed-entries patients was used. 13 common clinical parameters including numerical and categorical values were selected. Missing values in a missing-at-random (MAR) and missing-completely-at-random (MCAR) fashion from 0% to 60% were simulated and consequently imputed using the mean, hot-deck, multiple imputation by chained equations, expectation maximization method and listwise deletion. The performance was assessed by the root mean squared error, the absolute bias and the performance of a linear model for discharge mRS prediction.\n\nResultsListwise deletion was the worst performing method and started to be significantly worse than any imputation method from 2% (MAR) and 3% (MCAR) missing values on. The underlying missing value mechanism seemed to have a crucial influence on the identified best performing imputation method. Consequently no single imputation method outperformed all others. A significant performance drop of the linear model started from 11% (MAR+MCAR) and 18% (MCAR) missing values.\n\nConclusionsIn the presented case study of a typical clinical stroke dataset we confirmed that listwise deletion should be avoided for dealing with missing values. Our findings indicate that the underlying missing value mechanism and other dataset characteristics strongly influence the best choice of imputation method. For future studies with similar data structure, we thus suggest to use the developed framework in this study to select the most suitable imputation method for a given dataset prior to analysis.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A machine learning approach to identifying important features for achieving step thresholds in individuals with chronic stroke 95%
- Identification of high-risk COVID-19 patients using machine learning 94%
- Leveraging Machine Learning for Enhanced and Interpretable Risk Prediction of Venous Thromboembolism in Acute Ischemic Stroke Care 94%
Similar papers in this journal
- Deep learning approach for automatic assessment of schizophrenia and bipolar disorder in patients using R-R intervals 92%
- Uncertainty quantification in cerebral circulation simulations focusing on the collateral flow: Surrogate model approach with machine learning 91%
- A new machine learning method for cancer mutation analysis 91%
Similar papers in this journal
- Machine learning for classifying chronic kidney disease and predicting creatinine levels using at-home measurements 94%
- Predicting Car Accident Severity in Northwest Ethiopia: A Machine Learning Approach Leveraging Driver, Environmental, and Road Conditions 92%
- Selecting the most important self-assessed features for predicting conversion to Mild Cognitive Impairment with Random Forest and Permutation-based methods 92%
Similar papers in this journal
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 96%
- On the predictability of postoperative complications for cancer patients: a Portuguese cohort study 94%
- Combining symbolic regression with the Cox proportional hazards model improves prediction of heart failure deaths 94%
Similar papers in this journal
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 92%
- Learning from local to global - an efficient distributed algorithm for modeling time-to-event data 92%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.