Comparison of Imputation Strategies for Incomplete Electronic Health Data
Zhang, S.; Zhang, Z.; Hong, S.; Liu, H.; Zhou, y.
Show abstract
Missing data is a persistent challenge in electronic health records (EHRs), often compromising data integrity and limiting the effectiveness of predictive models in healthcare. This study systematically evaluates five widely used imputation strategies--GAIN, MICE, Median, MissForest, and MIWAE--across three real-world clinical datasets under varying missingness mechanisms (MCAR, MAR, and MNAR) and missingness rates (10%-90%). We assessed imputation quality using multiple statistical measures and examined the relationship between imputation accuracy and downstream classification performance. Our results show that MICE and MissForest consistently outperform other methods across most scenarios, while deep learning-based approaches such as GAIN exhibit high instability under MAR and MNAR, particularly at higher missingness levels. Furthermore, imputation quality does not always align with classification performance, underscoring the need to consider task-specific goals when selecting imputation strategies. We also provide a practical framework summarizing method recommendations based on missingness type and rate, aiming to support robust data preprocessing decisions in clinical AI applications.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Actionability of Synthetic Data in a Heterogeneous and Rare Healthcare Demographic; Adolescents and Young Adults (AYAs) with Cancer 93%
- Towards Predicting 30-Day Readmission among Oncology Patients: Identifying Timely and Actionable Risk Factors 92%
- Using Adversarial Images to Assess the Stability of Deep Learning Models Trained on Diagnostic Images in Oncology 91%
Similar papers in this journal
- Synthetic data for privacy-preserving clinical risk prediction 94%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 94%
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.