Mission imputable: Effects of missing data processing on infectious disease detection and prognosis
Roy, S. S.; Nguyen, N. T.; Zuniga, A.; Sarhaddi, F.; Lagerspetz, E.; Flores, H.; Nurmi, P.
Show abstract
BackgroundMissing data in medical datasets poses significant challenges for developing effective AI/ML pipelines. Inaccurate imputation can lead to biased results, reduced model performance, and compromised clinical insights. Understanding how different imputation methods affect AI/ML model performance is crucial for ensuring accurate clinical findings. ObjectiveThis study systematically investigates the effects of different imputation methods on AI/ML model performance and the clinical implications of these methods. MethodsWe investigate the impact of four different missing data strategies on the performance of common classification algorithms for analyzing medical data. The performance was evaluated based on sensitivity and specificity metrics for the tasks of predicting COVID-19 diagnosis and patient deterioration. We also perform feature analysis to understand the clinical implications the choice of imputation method has. ResultsOur findings show that the choice of imputation method significantly affects the performance of AI/ML techniques and the clinical conclusions drawn from the data. The optimal handling of missing values depends on (i) the composition of the features with missing values, (ii) the rate of missing values, and (iii) the pattern of the missing features. Using COVID-19 diagnosis and patient deterioration as representative examples of clinical tasks, our results indicate that MICE imputation yields the best overall performance, resulting in a 26% improvement in accuracy compared to baseline methods. Specifically, for predicting COVID-19 diagnosis, we achieved a sensitivity of 81% and specificity of 98%, while for patient deterioration, the sensitivity was 65% and specificity was 99%. ConclusionThis study demonstrates the critical impact of missing data imputation on AI/ML model performance and the clinical insights derived from these models. Our findings underscore the importance of selecting appropriate imputation techniques tailored to the specific characteristics of medical data to ensure accurate and reliable AI/ML predictions.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Prediction of Sepsis Mortality in ICU Patients Using Machine Learning Methods 97%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 96%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 95%
Similar papers in this journal
- Modeling physician variability to prioritize relevant medical record information 97%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 95%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 93%
Similar papers in this journal
Similar papers in this journal
- Improving irregular temporal modeling by integrating synthetic data to the electronic medical record using conditional GANs: a case study of fluid overload prediction in the intensive care unit 96%
- AI-MET: A Deep Learning-based Clinical Decision Support System for Distinguishing Multisystem Inflammatory Syndrome in Children from Endemic Typhus 96%
- Machine Learning Interpretability Methods to Characterize the Importance of Hematologic Biomarkers in Prognosticating Patients with Suspected Infection 96%
Similar papers in this journal
- Machine learning for classifying chronic kidney disease and predicting creatinine levels using at-home measurements 96%
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 96%
- Machine learning to predict retention and viral suppression in South African HIV treatment cohorts 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.