The Art of Not Knowing: Accommodating Structured Missingness in Biomedical Research
Rehak Buckova, B.; Fraza, C.; Koldbaek, C.; Barthel Flaaten, C.; Sofie Saether, L.; Cattarinussi, G.; Sando Ambrosen, K. M.; Andreassen, O. A.; Westlye, L. T.; Beckmann, C.; Ebdrup, B.; Dazzan, P.; Ueland, T.; Mitra, R.; Marquand, A.
Show abstract
Missing data remain a ubiquitous and critical challenge in large-scale clinical studies. Despite advances in imputation, most existing methods fail to address structured missingness, where data are missing according a deterministic pattern and which arise due to systematic patterns introduced by experimental design, site protocols, or cohort differences. These patterns violate key assumptions of most imputation algorithms, yet their impact is rarely evaluated. We demonstrate that structured missingness is a fundamental challenge to drawing valid inferences from standard imputation techniques. First, we present a comprehensive framework for understanding and accommodating the effects of structured missingness. Next, we show through simulations and real-world psychometric data with structured missingness, that widely used algorithms optimised for numerical precision (e.g., Extra Trees, AutoComplete) underperform due to site effects, while donor-based methods (e.g., MICE, hierarchical MICE) better preserve multivariate structure. We propose a novel hierarchical approach that provides optimal performance in simulated and experimental data. Finally, we show that commonly used accuracy metrics, such as mean squared error can obscure these failures, and are therefore inadequate for the evaluation of structured missingness. In contrast, other divergence-based metrics offer a more sensitive and interpretable alternative. We apply this approach to harmonising psychometric data across cohorts, which provides excellent item-level alignment across different instruments. Our study highlights the need for a paradigm shift in handling missing data within biomedical research, moving beyond conventional imputation frameworks to develop tools that can account for structured missingness. This shift is essential for ensuring reliable inference in multi-site clinical studies, precision medicine, and large-scale population analyses.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enabling interpretable machine learning for biological data with reliability scores 93%
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 93%
- Using random forests to uncover the predictive power of distance-varying cell interactions in tumor microenvironments 92%
Similar papers in this journal
Similar papers in this journal
- Compressive Big Data Analytics: An Ensemble Meta-Algorithm for High-dimensional Multisource Datasets 95%
- ACMTF-R: supervised multi-omics data integration uncovering shared and distinct outcome-associated variation 93%
- Analyzing Biomarker Discovery: Estimating the Reproducibility of Biomarker Sets 92%
Similar papers in this journal
- A Stability-Enhanced Lasso Approach for Covariate Selection in Non-Linear Mixed Effect Model 94%
- Penalized reduced rank regression for multi-outcome survival data supports a common metabolic risk score for age-related diseases 93%
- Two-phase sample selection strategies for design and analysis in post-genome wide association fine-mapping studies 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.