Back

Privacy-Preserving Matching for Federated Causal Inference in Multicentre Patient Cohorts

Gusinow, R.; Morgan, A. S.; Canziani, L. M.; Zeitlin, J.; Kim, M.; Gentilotti, E.; Ghosn, J.; Florence, A.-M.; Tami, A.; Toschi, A.; Palacios-Baena, Z. R.; Tacconelli, E.; Hasenauer, J.

2026-07-19 epidemiology
10.64898/2026.07.16.26358171 medRxiv
Show abstract

Causal effect estimates can often be biased in clinical and epidemiological studies as patient cohorts frequently exhibit substantial covariate imbalances between treated and control groups, often amplified in multicentre studies due to heterogeneous recruitment, clinical practice, and case mix. Covariate balancing methods are therefore essential for valid causal inference. However, their application becomes challenging when data are distributed across cohorts and cannot be pooled because of privacy, legal, or institutional constraints, leaving a gap in practical methods for causal effect estimation in federated and imbalanced clinical data settings. We develop a privacy-preserving framework for covariate balancing and causal effect estimation across distributed data providers, combining federated aggregation with differential privacy to enable propensity score subclassification and matching without sharing individual-level records. Matching relies on non-disclosive quantities and differentially private distance evaluation, and the resulting matched subsets remain local to each server. Balance can be assessed through federated diagnostics and privacy-preserving visualisations, and we provide secure estimators for average treatment effects with associated uncertainty quantification. We implement this framework in the DataSHIELD federated analysis platform via 2 R packages. In simulations, we demonstrate agreement between federated and centralised analyses in the absence of privacy noise and quantify the bias--variance trade-offs induced by differential privacy. We illustrate applicability in two multinational settings-a Long COVID cohort and very preterm birth cohorts-showing that the approach enables practical causal analyses under real-world data protection constraints. The DataSHIELD packages are available on Github. Additional methodological details are provided in the Supplementary Material.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 8%
18.6%
2
Nature Methods
385 papers in training set
Top 2%
5.5%
3
Nature Medicine
125 papers in training set
Top 0.3%
5.5%
4
Patterns
78 papers in training set
Top 0.3%
4.9%
5
Nature Genetics
286 papers in training set
Top 1%
4.9%
6
International Journal of Epidemiology
88 papers in training set
Top 0.3%
4.3%
7
eLife
5828 papers in training set
Top 37%
2.8%
8
Science Advances
1243 papers in training set
Top 12%
2.8%
9
npj Digital Medicine
118 papers in training set
Top 2%
2.8%
50% of probability mass above
10
Nature
645 papers in training set
Top 5%
2.5%
11
Nature Biotechnology
172 papers in training set
Top 2%
2.4%
12
BMC Medical Research Methodology
47 papers in training set
Top 0.5%
2.1%
13
American Journal of Epidemiology
67 papers in training set
Top 0.6%
1.9%
14
Statistics in Medicine
40 papers in training set
Top 0.3%
1.9%
15
Cell Reports Methods
165 papers in training set
Top 1%
1.9%
16
PLOS ONE
5266 papers in training set
Top 46%
1.9%
17
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 26%
1.9%
18
PLOS Computational Biology
1863 papers in training set
Top 14%
1.7%
19
Scientific Reports
3612 papers in training set
Top 55%
1.7%
20
Science Translational Medicine
127 papers in training set
Top 2%
1.5%
21
Genome Research
468 papers in training set
Top 4%
1.4%
22
Wellcome Open Research
67 papers in training set
Top 0.8%
1.4%
23
Cell Systems
201 papers in training set
Top 3%
1.4%
24
Bioinformatics
1204 papers in training set
Top 7%
1.3%
25
Epidemics
116 papers in training set
Top 1%
1.3%
26
The American Journal of Human Genetics
234 papers in training set
Top 2%
1.1%
27
Biometrics
23 papers in training set
Top 0.3%
1.0%
28
Communications Medicine
113 papers in training set
Top 4%
1.0%
29
Journal of Clinical Epidemiology
31 papers in training set
Top 0.7%
0.9%
30
European Journal of Epidemiology
43 papers in training set
Top 0.7%
0.8%