Cross-site imputation for recovering variables without individual pooled data
Thiesmeier, R.; Madley-Dowd, P.; Orsini, N.; Ahlqvist, V.
Show abstract
In multi-site studies, it is common for some sites not to have recorded key variables. Although it is theoretically possible to use data from sites with recorded observations to impute the missing values, this process becomes challenging when data pooling is not feasible due to logistic or legal constraints. We therefore propose a multiple imputation approach -- cross-site imputation -- to recover any variables across sites without the need to pool individual-level data. The solution involves transporting predicted regression coefficients and variances from studies with observed data to impute missing variables at sites without data. The approach is illustrated in an applied example of recovering confounders across Swedish hospitals, and theoretical considerations are outlined. Given the increasing importance of multi-site studies in observational research, cross-site imputation could offer a practical approach for imputing variables that have not been recorded in some study sites.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 95%
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 93%
- Comparison of Bayesian networks, G-estimation and linear models to estimate causal treatment effects in aggregated N-of-1 trials 91%
Similar papers in this journal
- Common misconceptions held by health researchers when interpreting linear regression assumptions, a cross-sectional study 94%
- Introducing riskCommunicator: an R package to obtain interpretable effect estimates for public health 93%
- Wellbeing Impact Study of High-Speed 2 (WISH2): Protocol for a mixed-methods examination of the impact of major transport infrastructure development on mental health and wellbeing 93%
Similar papers in this journal
Similar papers in this journal
- Causal Forests versus Inverse Probability of Treatment Weighting to adjust for Cluster-Level Confounding: A Parametric and Plasmode Simulation Study based on US Hosptial Electronic Health Record Data 93%
- Bias amplification of unobserved confounding in pharmacoepidemiological studies using indication-based sampling: there is no free lunch in restricting the sample to those with a particular drug-indication 92%
- INSIGHT: A Tool for Fit-for-Purpose Evaluation and Quality Assessment of Observational Data Sources for Real World Evidence on Medicine and Vaccine Safety 92%
Similar papers in this journal
- Toward Evaluation of Disseminated Effects of Medications for Opioid Use Disorder within Provider-Based Clusters Using Routinely-Collected Health Data 93%
- Power Analysis for Stepped Wedge Trials with Two Treatments 92%
- Sensitivity to missing not at random dropout in clinical trials: use and interpretation of the Trimmed Means Estimator 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.