Multiple imputation of missing data under missing at random: compatible imputation models are not sufficient to avoid bias
Curnow, E.; Carpenter, J. R.; Heron, J. E.; Cornish, R. P.; Rach, S.; Didelez, V.; Langeheine, M.; Tilling, K.
Show abstract
BackgroundEpidemiological studies often have missing data. Multiple imputation (MI) is a commonly-used strategy for such studies. MI guidelines for structuring the imputation model have focused on compatibility with the analysis model, but not on the need for the (compatible) imputation model(s) to be correctly specified. Standard (default) MI procedures use simple linear functions. We examine the bias this causes and performance of methods to identify problematic imputation models, providing practical guidance for researchers. MethodsBy simulation and real data analysis, we investigated how imputation model mis-specification affected MI performance, comparing results with complete records analysis (CRA). We considered scenarios in which imputation model mis-specification occurred because (i) the analysis model was mis-specified, or (ii) the relationship between exposure and confounder was mis-specified. ResultsMis-specification of the relationship between outcome and exposure, or between exposure and confounder in the imputation model for the exposure, could result in substantial bias in CRA and MI estimates (in addition to any bias in the full-data estimate due to analysis model mis-specification). MI by predictive mean matching could mitigate for model mis-specification. Model mis-specification tests were effective in identifying mis-specified relationships. These could be easily applied in any setting in which CRA was, in principle, valid and data were missing at random (MAR). ConclusionWhen using MI methods that assume data are MAR, compatibility between the analysis and imputation models is necessary, but is not sufficient to avoid bias. We propose an easy-to-follow, step-by-step procedure for identifying and correcting mis-specification of imputation models.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 96%
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 96%
- Comparison of Bayesian networks, G-estimation and linear models to estimate causal treatment effects in aggregated N-of-1 trials 95%
Similar papers in this journal
- Sensitivity to missing not at random dropout in clinical trials: use and interpretation of the Trimmed Means Estimator 96%
- A Double Machine Learning Approach for the Evaluation of COVID-19 Vaccine Effectiveness under the Test-Negative Design: Analysis of Québec Administrative Data 94%
- Logistic mixed-effects model analysis with pseudo-observations for estimating risk ratios in clustered binary data analysis 94%
Similar papers in this journal
- Introducing riskCommunicator: an R package to obtain interpretable effect estimates for public health 95%
- Common misconceptions held by health researchers when interpreting linear regression assumptions, a cross-sectional study 93%
- Using numerical modelling and simulation to assess the ethical burden in clinical trials and how it relates to the proportion of responders in a trial sample 93%
Similar papers in this journal
- Estimating and Testing an Index of Bias Attributable to Composite Outcomes in Comparative Studies 94%
- Quantitative bias analysis methods for summary level epidemiologic data in the peer-reviewed literature: a systematic review 93%
- Diagnostic test accuracy in longitudinal study settings: Theoretical approaches with use cases from clinical practice 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.