ImputeBench: Benchmarking Single Imputation Methods
Richter, R.; Tavares, J. F.; Miloschewski, A.; Breteler, M. M. B.; Mukherjee, S.
Show abstract
Biomedical data often contain missing values and in many applications missing value imputation (MVI) is an important part of the data analysis work-flow. However, the performance of MVI methods depends on details of the joint distribution of data and missingness patterns that are typically unknown in practice, making an a priori choice of MVI method challenging. Furthermore, technical assumptions underlying MVI methods can be hard to directly verify in practice. Motivated by these issues, in this paper, we propose an approach for the context-specific selection of MVI methods. Due to the fact that different methods may work well in different cases we argue for a move away from a "one size fits all" view and put forward in this paper a standardized, empirical approach in which MVI methods are benchmarked in the specific context of a problem of interest. We connect our work to the large body of MVI research, along the way refining definitions of missing at random and missing not at random and providing a detailed review of existing work on benchmarking. Our approach can be tailored to reflect specific assumptions on missingness patterns, allowing for application in diverse applied problems. Furthermore, in addition to using real data, we study benchmarking via data simulation spanning a broad range of properties, such as latent factors, non-linearity and multi-modality, with interpretable simulation parameters that are amenable to user specification. The approaches we propose can be used to (i) select an MVI method for a given data set or (ii) benchmark a novel MVI method across a range of regimes. Alongside the general protocol, we provide a specific, reproducible implementation (in the R-package ImputeBench, available under github.com/richterrob/ImputeBench) that gives users a ready-to-use tool for MVI selection and assessment. We illustrate the use of ImputeBench to study the behaviour of a range of existing imputation methods (k-nn, soft impute, missForest, MICE) in the context of real data from an ongoing large-scale population-level study.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Penalized reduced rank regression for multi-outcome survival data supports a common metabolic risk score for age-related diseases 95%
- A Stability-Enhanced Lasso Approach for Covariate Selection in Non-Linear Mixed Effect Model 93%
- Assessing Reproducibility of High-throughput Experiments in the Case of Missing Data. 93%
Similar papers in this journal
- GPerturb: Gaussian process modelling of single-cell perturbation data 93%
- Statistical modeling, estimation, and remediation of sample index hopping in multiplexed droplet-based single-cell RNA-seq data 92%
- Permutation-based Identification of Important Biomarkers for Complex Diseases via Black-box Models 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.