Back

Simulation of synthetic health records for assessment of causal inference methods for vaccine efficacy

Velasco Pardo, V.; Daines, L.; Katikireddi, S. V.; Ritchie, L.; Robertson, C.; Simpson, C. R.; McCowan, C.; Swallow, B.

2026-07-19 infectious diseases
10.64898/2026.07.17.26358308 medRxiv
Show abstract

Background During the COVID-19 pandemic, public health agencies used near real-time observational data to answer questions regarding vaccine effectiveness. However, traditional observational methods do not allow conclusions regarding counterfactual scenarios to be drawn from clinical data. Counterfactuals, which are outcomes that would have occurred under alternative interventions, can be used to formally assess the causal effects of public health interventions on health outcomes while accounting for the effects of confounding. Ideally individual patient data is used for the development of counterfactuals. Low-fidelity synthetic data may be useful for advancing methodological development where governance and privacy constraints prohibit access to sensitive personal data. Methods We simulated synthetic datasets based on the EAVE-II COVID-19 platform which has been limited to use for surveillance purposes. EAVE-II includes almost all resident people in Scotland registered with qualified general medical practitioners. Patient characteristics were simulated to reflect the known distribution of the Scottish population, accounting for dependencies between variables. Each synthetic dataset was encoded to different realistic scenarios for EAVEII 'ground truth' vaccine rollout and effectiveness results, explicitly stating the causal and confounding mechanisms, using a statistically sound method based on marginal structural models. Synthetic datasets of 100,000 individuals were then generated across five confounding scenarios and five severe outcome types. Results In scenarios with weak confounding, both unweighted and inverse probability of treatment weighted (IPTW) logistic regression recovered the true causal parameters. As confounding strength increased, only weighted models recovered the true mechanism. Conclusions Low-fidelity synthetic datasets simulated from EAVE-II data analysts to build and test causal inference pipelines, develop novel analysis pipelines, and train new researchers while awaiting access to real data. We showed how to generate synthetic datasets from a marginal structural model under different confounding scenarios.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
American Journal of Epidemiology
67 papers in training set
Top 0.1%
26.4%
2
BMC Medical Research Methodology
47 papers in training set
Top 0.1%
9.6%
3
Statistics in Medicine
40 papers in training set
Top 0.1%
6.2%
4
Epidemics
116 papers in training set
Top 0.4%
5.5%
5
International Journal of Epidemiology
88 papers in training set
Top 0.4%
3.2%
50% of probability mass above
6
PLOS ONE
5266 papers in training set
Top 38%
3.2%
7
PLOS Computational Biology
1863 papers in training set
Top 11%
2.6%
8
Medical Decision Making
12 papers in training set
Top 0.1%
2.4%
9
Epidemiology
32 papers in training set
Top 0.2%
2.4%
10
Journal of Clinical Epidemiology
31 papers in training set
Top 0.4%
1.9%
11
Nature Communications
5641 papers in training set
Top 43%
1.9%
12
Vaccine
203 papers in training set
Top 1%
1.7%
13
European Journal of Epidemiology
43 papers in training set
Top 0.4%
1.5%
14
BMJ
51 papers in training set
Top 0.6%
1.4%
15
BMJ Open
601 papers in training set
Top 10%
1.4%
16
Biometrics
23 papers in training set
Top 0.2%
1.4%
17
eLife
5828 papers in training set
Top 58%
1.1%
18
Trials
29 papers in training set
Top 0.8%
1.1%
19
Wellcome Open Research
67 papers in training set
Top 1%
1.1%
20
Clinical Trials
11 papers in training set
Top 0.3%
1.1%
21
Communications Medicine
113 papers in training set
Top 4%
1.1%
22
PLOS Global Public Health
344 papers in training set
Top 7%
1.0%
23
Biostatistics
24 papers in training set
Top 0.3%
1.0%
24
Pharmacoepidemiology and Drug Safety
18 papers in training set
Top 0.5%
0.8%
25
Value in Health
11 papers in training set
Top 0.3%
0.8%
26
Frontiers in Public Health
148 papers in training set
Top 6%
0.8%
27
JMIR Medical Informatics
18 papers in training set
Top 1%
0.6%
28
Nature Human Behaviour
95 papers in training set
Top 3%
0.6%
29
Patterns
78 papers in training set
Top 3%
0.6%
30
Science
477 papers in training set
Top 10%
0.6%