Multiple imputation of missing data under missing at random: including a collider as an auxiliary variable in the imputation model can induce bias
Curnow, E.; Tilling, K.; Heron, J.; Cornish, R. P.; Carpenter, J. R.
Show abstract
Epidemiological studies often have missing data, which are commonly handled by multiple imputation (MI). In MI, in addition to those required for the substantive analysis, imputation models often include other variables ("auxiliary variables"). Auxiliary variables that predict the partially observed variables can reduce the standard error (SE) of the MI estimator and, if they also predict the probability that data are missing, reduce bias due to data being missing not at random. However, guidance for choosing auxiliary variables is lacking. We examine the consequences of a poorly-chosen auxiliary variable: if it shares a common cause with the partially observed variable and the probability that it is missing (i.e. it is a "collider"), its inclusion can induce bias in the MI estimator and may increase SE. We quantify, both algebraically and by simulation, the magnitude of bias and SE when either the exposure or outcome are incomplete. When the substantive analysis outcome is partially observed, the bias can be substantial, relative to the magnitude of the exposure coefficient. In settings in which complete records analysis is valid, the bias is smaller when the exposure is partially observed. However, bias can be larger if the outcome also causes missingness in the exposure. When using MI, it is important to examine, through a combination of data exploration and considering plausible casual diagrams and missingness mechanisms, whether potential auxiliary variables are colliders. Contribution to the field statementIn multiple imputation (MI), in addition to those required for the substantive analysis, imputation models often include other variables ("auxiliary variables"). Auxiliary variables that predict the partially observed variables can reduce the standard error (SE) of the MI estimator and, if they also predict the probability that data are missing, reduce bias due to data being missing not at random. We examine the consequences of a poorly-chosen auxiliary variable: if it shares a common cause with the partially observed variable and the probability that it is missing (i.e. it is a "collider"), its inclusion can induce bias in the MI estimator and may increase SE. We demonstrate that when the substantive analysis outcome is partially observed, the bias can be substantial, relative to the magnitude of the exposure coefficient. In settings in which complete records analysis is valid, the bias is smaller when the exposure is partially observed. However, bias can be larger if the outcome also causes missingness in the exposure. We recommmend a combination of data exploration and consideration of plausible casual diagrams and missingness mechanisms to examine whether potential auxiliary variables are colliders.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Sensitivity to missing not at random dropout in clinical trials: use and interpretation of the Trimmed Means Estimator 95%
- A Double Machine Learning Approach for the Evaluation of COVID-19 Vaccine Effectiveness under the Test-Negative Design: Analysis of Québec Administrative Data 95%
- Bias reduction and inference for electronic health record data under selection and phenotype misclassification: three case studies 94%
Similar papers in this journal
- Negative Control Exposures: Causal effect Identifiability and Use in Probabilistic-Bias and Bayesian Analyses with Unmeasured Confounders 95%
- Sensitivity and Uncertainty Analysis for Two-Stream Capture-Recapture Methods in Disease Surveillance 93%
- Assessing Direct and Spillover Effects of Intervention Packages in Network-Randomized Studies 91%
Similar papers in this journal
- Causal Mediation Analysis with Multiple Causally Ordered and Non-ordered Mediators based on Summarized Genetic Data 94%
- Two-Stage Multivariate Mendelian Randomization on Multiple Outcomes with Mixed Distributions 94%
- A bivariate zero-inflated negative binomial model and its applications to biomedical settings 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.