Back

Impact of missing data correlated with labels to be predicted in neurodegeneration classification tasks

Prakash, M.; Tohka, J.

2025-01-24 bioinformatics
10.1101/2025.01.23.634117 bioRxiv
Show abstract

We introduce a new subtype of Missing Not at Random (MNAR) data, where the missingness is correlated with the labels (y) to be predicted, termed (y)-dependent MNAR. We demonstrate that this subtype can significantly bias the estimation of performance metrics in typical machine learning tasks. Unbiased error estimation is crucial in predictive modeling to accurately assess model performance, identify potential biases, and ensure generalizability to new, unseen data. We explore the effects of imputing this new subtype of MNAR and compare it with general missing types, namely Missing at Random (MAR) and Missing Completely at Random (MCAR). Our comparison analysis employs both synthetic and clinical datasets, including the Alzheimers Disease Neuroimaging Initiative (ADNI) dataset, the Parkinsons Progression Markers Initiative (PPMI) dataset, and the Anti-Amyloid Treatment in Asymptomatic Alzheimers Disease (A4) dataset. After introducing missingness into the datasets, we trained different classifiers paired with various imputation methods and measured repeated cross-validation test metrics. Our findings reveal that datasets with non-ignorable missing types (MNAR) exhibit a strong bias compared to ignorable types (MAR and MCAR) in downstream analysis. Non-linear classifiers tend to exploit patterns from imputed data, particularly when the imputed values correlate with the target label (y), which can lead to unreliable estimation of the generalization error. Mean and median imputations proved to be more robust than tree-based or gradient boosting methods.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.