Impact of missing data correlated with labels to be predicted in neurodegeneration classification tasks
Prakash, M.; Tohka, J.
Show abstract
We introduce a new subtype of Missing Not at Random (MNAR) data, where the missingness is correlated with the labels (y) to be predicted, termed (y)-dependent MNAR. We demonstrate that this subtype can significantly bias the estimation of performance metrics in typical machine learning tasks. Unbiased error estimation is crucial in predictive modeling to accurately assess model performance, identify potential biases, and ensure generalizability to new, unseen data. We explore the effects of imputing this new subtype of MNAR and compare it with general missing types, namely Missing at Random (MAR) and Missing Completely at Random (MCAR). Our comparison analysis employs both synthetic and clinical datasets, including the Alzheimers Disease Neuroimaging Initiative (ADNI) dataset, the Parkinsons Progression Markers Initiative (PPMI) dataset, and the Anti-Amyloid Treatment in Asymptomatic Alzheimers Disease (A4) dataset. After introducing missingness into the datasets, we trained different classifiers paired with various imputation methods and measured repeated cross-validation test metrics. Our findings reveal that datasets with non-ignorable missing types (MNAR) exhibit a strong bias compared to ignorable types (MAR and MCAR) in downstream analysis. Non-linear classifiers tend to exploit patterns from imputed data, particularly when the imputed values correlate with the target label (y), which can lead to unreliable estimation of the generalization error. Mean and median imputations proved to be more robust than tree-based or gradient boosting methods.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Compressive Big Data Analytics: An Ensemble Meta-Algorithm for High-dimensional Multisource Datasets 96%
- Random forest model for feature-based Alzheimer's disease conversion prediction from early mild cognitive impairment subjects 95%
- c-Triadem: A constrained, explainable deep learning model to identify novel biomarkers in Alzheimer’s disease 94%
Similar papers in this journal
- Selecting the most important self-assessed features for predicting conversion to Mild Cognitive Impairment with Random Forest and Permutation-based methods 95%
- A novel interpretable deep transfer learning combining diverse learnable parameters for improved T2D prediction based on single-cell gene regulatory networks 93%
- Predicting cognitive decline in a low-dimensional representation of brain morphology 93%
Similar papers in this journal
- BenchXAI: Comprehensive Benchmarking of Post-hoc Explainable AI Methods on Multi-Modal Biomedical Data 96%
- Alzheimer Disease Knowledge Graph Enhances Knowledge Discovery and Disease Prediction 93%
- Deep Multimodal Graph-Based Network for Survival Prediction from Highly Multiplexed Images and Patient Variables 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.