Simulation of synthetic health records for assessment of causal inference methods for vaccine efficacy
Velasco Pardo, V.; Daines, L.; Katikireddi, S. V.; Ritchie, L.; Robertson, C.; Simpson, C. R.; McCowan, C.; Swallow, B.
Show abstract
Background During the COVID-19 pandemic, public health agencies used near real-time observational data to answer questions regarding vaccine effectiveness. However, traditional observational methods do not allow conclusions regarding counterfactual scenarios to be drawn from clinical data. Counterfactuals, which are outcomes that would have occurred under alternative interventions, can be used to formally assess the causal effects of public health interventions on health outcomes while accounting for the effects of confounding. Ideally individual patient data is used for the development of counterfactuals. Low-fidelity synthetic data may be useful for advancing methodological development where governance and privacy constraints prohibit access to sensitive personal data. Methods We simulated synthetic datasets based on the EAVE-II COVID-19 platform which has been limited to use for surveillance purposes. EAVE-II includes almost all resident people in Scotland registered with qualified general medical practitioners. Patient characteristics were simulated to reflect the known distribution of the Scottish population, accounting for dependencies between variables. Each synthetic dataset was encoded to different realistic scenarios for EAVEII 'ground truth' vaccine rollout and effectiveness results, explicitly stating the causal and confounding mechanisms, using a statistically sound method based on marginal structural models. Synthetic datasets of 100,000 individuals were then generated across five confounding scenarios and five severe outcome types. Results In scenarios with weak confounding, both unweighted and inverse probability of treatment weighted (IPTW) logistic regression recovered the true causal parameters. As confounding strength increased, only weighted models recovered the true mechanism. Conclusions Low-fidelity synthetic datasets simulated from EAVE-II data analysts to build and test causal inference pipelines, develop novel analysis pipelines, and train new researchers while awaiting access to real data. We showed how to generate synthetic datasets from a marginal structural model under different confounding scenarios.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Double Machine Learning Approach for the Evaluation of COVID-19 Vaccine Effectiveness under the Test-Negative Design: Analysis of Québec Administrative Data 95%
- Estimation of Vaccine Efficacy for Variants that Emerge After the Placebo Group Is Vaccinated 93%
- Assessing Covariate Balance with Small Sample Sizes 93%
Similar papers in this journal
- Potential Biases in Test-Negative Design Studies of COVID-19 Vaccine Effectiveness Arising from the Inclusion of Asymptomatic Individuals 94%
- Misclassification of yellow fever vaccination status revealed through hierarchical Bayesian modeling 92%
- Analyses using multiple imputation need to consider missing data in auxiliary variables 92%
Similar papers in this journal
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 94%
- Comparison of Bayesian networks, G-estimation and linear models to estimate causal treatment effects in aggregated N-of-1 trials 93%
- Over- and under-estimation of vaccine effectiveness 92%
Similar papers in this journal
- Incorporating data from multiple endpoints in the analysis of clinical trials: example from RSV vaccines 94%
- Assessing Direct and Spillover Effects of Intervention Packages in Network-Randomized Studies 93%
- Use of recently vaccinated individuals to detect bias in test-negative case-control studies of COVID-19 vaccine effectiveness 93%
Similar papers in this journal
- The methodologies to assess the effects of non-pharmaceutical interventions during COVID-19: a systematic review 91%
- Mendelian randomisation for mediation analysis: current methods and challenges for implementation 91%
- A comparison of regression discontinuity and propensity score matching to estimate the causal effects of statins using electronic health records 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.