Selection bias due to omitting interactions from inverse probability weighting
Wen, L.; Tilling, K.; Cornish, R. P.; Hughes, R.; Gkatzionis, A.
Show abstract
Inverse probability weighting (IPW) is often used to adjust for selection bias, typically using a simple logit model without interactions as a missingness model. However, the size of the selection bias depends on the interaction between exposure and outcome in their effect on selection - implying that it may be important to include interactions in the IPW model. Via a simulation and a real-data application we compare the performance of IPW with and without interaction terms to estimate a regression coefficient. The simulation study shows that IPW including interactions gives less biased estimates than IPW without interactions in all scenarios studied. Importantly, IPW using a logit model with no interactions often gives estimates close to the complete case analysis (CCA) - perhaps giving false reassurance that results are robust to selection bias. The real-data application investigates the association between unemployment and sleep duration, using data from Understanding Society. IPW including interactions suggests that unemployment is associated with a reduction in sleep duration of around 23 (9, 38) minutes, compared to 27 (14, 40) minutes for IPW without interactions, and 31 (19, 43) minutes for CCA. We strongly recommend including interactions in missingness models to adjust for selection bias.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Reweighting the UK Biobank to reflect its underlying sampling population substantially reduces pervasive selection bias due to volunteering 95%
- Bias in two-sample Mendelian randomization when using heritable covariable-adjusted summary associations 93%
- Estimation of time-varying causal effects with multivariable Mendelian randomization: some cautionary notes 93%
Similar papers in this journal
- Analyses using multiple imputation need to consider missing data in auxiliary variables 96%
- Combining longitudinal data from different cohorts to examine the life-course trajectory 91%
- A Hierarchical Approach Using Marginal Summary Statistics for Multiple Intermediates in a Mendelian Randomization or Transcriptome Analysis 91%
Similar papers in this journal
- Mendelian randomisation for mediation analysis: current methods and challenges for implementation 95%
- How to mitigate selection bias in COVID-19 surveys: evidence from five national cohorts 93%
- The effect of long-term adherence to physical activity recommendations in midlife on plasma proteins associated with frailty in the Atherosclerosis Risk in Communities (ARIC) study 91%
Similar papers in this journal
- Quantitative bias analysis in practice: Review of software for regression with unmeasured confounding 95%
- Covid-19 prevalence estimation by random sampling in the wider population - Optimal sample pooling under varying assumptions about true prevalence 92%
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 92%
Similar papers in this journal
- Negative Control Exposures: Causal effect Identifiability and Use in Probabilistic-Bias and Bayesian Analyses with Unmeasured Confounders 95%
- The Contribution of Health Behaviors to Depression Risk across Birth Cohorts 93%
- Assessing Direct and Spillover Effects of Intervention Packages in Network-Randomized Studies 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.