Propensity-score matching with GAN-generated observations from electronic health records: simulation study and application to the evaluation of prone positioning in COVID-19 patients under mechanical ventilation
Bouvarel, B.; Glemain, B.; Carrat, F.; Lapidus, N.
Show abstract
BackgroundPropensity score (PS) methods are widely used in observational studies to estimate causal effects, but they often exclude patients due to a lack of comparable counterparts, leading to reduced power and potential bias. Generative adversarial networks (GANs) have shown promise in creating synthetic data, but their application to causal inference remains underexplored. Synthetic data could be used as plausible counterfactuals, potentially mitigating the issues of the PS methods. This study evaluates the integration of GAN-generated synthetic observations into propensity score matching (PSM) to improve the emulation of RCTs, using both simulated and real-world electronic health record (EHR) data. MethodsA simulation study was conducted using with predefined confounding structures to compare traditional PSM against two hybrid approaches incorporating GAN-generated synthetic patients to partially or fully match the original sample of patients. Treatment effects were estimated via logistic regression, and performance was assessed by bias, standard error, alpha risk, power, and confidence interval coverage. The methods were applied to a real-world dataset of mechanically ventilated COVID-19 patients to evaluate the impact of early prone positioning on 28-day mortality. ResultsIn simulations, GAN-generated patients permitted to match all patients in the original sample, whereas PSM dropped up to 60% of them. While synthetic augmentation improved sample size, unadjusted use of synthetic matches led to underestimated standard errors and inflated type I error. Down-weighting matched synthetic data improved error control but did not consistently outperform PSM in bias or power. In the real-world application (n=1399), treatment effect estimates for prone positioning were similar across all methods and did not reach statistical significance. ConclusionGAN-augmented propensity score matching can reduce sample loss. However, its current application in causal inference through PS matching remains limited. Synthetic data do not contribute independent information and must be cautiously integrated to avoid misleading precision. While promising, current GAN implementations require methodological refinements before routine use in causal inference.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 96%
- Comparison of Bayesian networks, G-estimation and linear models to estimate causal treatment effects in aggregated N-of-1 trials 94%
- KMSubtraction: Reconstruction of unreported subgroup survival data utilizing published Kaplan-Meier survival curves 93%
Similar papers in this journal
- Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies 95%
- Automated Interpretable Discovery of Heterogeneous Treatment Effectiveness: A Covid-19 Case Study 95%
- SurvMaximin: Robust Federated Approach to Transporting Survival Risk Prediction Models 94%
Similar papers in this journal
- Learning from local to global - an efficient distributed algorithm for modeling time-to-event data 94%
- Analysis of Eligibility Criteria Clusters Based on Large Language Models for Clinical Trial Design 93%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 93%
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 92%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 92%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 92%
Similar papers in this journal
- Protocol for an observational study evaluating new approaches to modelling diagnostic information from large administrative hospital datasets 93%
- Estimating and Testing an Index of Bias Attributable to Composite Outcomes in Comparative Studies 92%
- Diagnostic test accuracy in longitudinal study settings: Theoretical approaches with use cases from clinical practice 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.