Double Machine Learning for Causal Inference in High-Dimensional Electronic Health Records
Du, M.; Guo, Y.; Li, X.; Catala, M.; Prieto-Alhambra, D.
Show abstract
BackgroundEstimating causal effects in observational health data is challenging due to confounding by indication. Traditional approaches such as inverse probability of treatment weighting (IPTW) rely on correct model specification, which is difficult in high-dimensional settings. We implemented an offset-based double machine learning (Offset-DML) practical framework for estimating binary treatment effects on the log-odds scale using logistic regression. MethodsWe have conducted a plasmode simulation study based on real-world clinical data, varying sample sizes (5,000, 10,000, 20,000) and outcome prevalence (5%, 10%, 20%) with 200 repetitions. We compared the performance of IPTW, stabilised IPTW, offset-DML (with and without cross-fitting), and high-dimensional DML (HD-DML). We measured and compared the performance of the different models with the following metrics: absolute bias, empirical standard error, and root mean square error relative to the true average causal effect. ResultsAcross most scenarios, DML-based approaches outperformed IPTW methods in terms of bias and empirical standard error, particularly in larger sample sizes. Offset-DML showed comparable performance to HD-DML while avoiding convergence issues observed with HD-DML in sparse data settings. All DML methods had overlapping confidence intervals in most scenarios. ConclusionOffset-DML is a practical and robust alternative for causal inference in high-dimensional health data. Future work should investigate extensions to other outcomes and diagnostics to assess confounding control. Key messagesO_LIDouble machine learning based methods consistently outperform IPTW regarding bias and empirical standard error, particularly in large sample sizes and sparse-data scenarios. C_LIO_LIOffset Double machine learning is a practical and robust binary causal effect estimation method in high-dimensional settings. C_LIO_LIUnlike high-dimensional Double machine learning, the offset-based Double machine learning approach demonstrated consistent convergence across all scenarios, including those with low outcome prevalence and small sample sizes. C_LI
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Double Machine Learning Approach for the Evaluation of COVID-19 Vaccine Effectiveness under the Test-Negative Design: Analysis of Québec Administrative Data 94%
- Sensitivity to missing not at random dropout in clinical trials: use and interpretation of the Trimmed Means Estimator 94%
- A Stability-Enhanced Lasso Approach for Covariate Selection in Non-Linear Mixed Effect Model 94%
Similar papers in this journal
- Potential Biases in Test-Negative Design Studies of COVID-19 Vaccine Effectiveness Arising from the Inclusion of Asymptomatic Individuals 93%
- Analyses using multiple imputation need to consider missing data in auxiliary variables 92%
- Obtaining prevalence estimates of COVID-19: A model to inform decision-making 92%
Similar papers in this journal
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 95%
- Comparison of Bayesian networks, G-estimation and linear models to estimate causal treatment effects in aggregated N-of-1 trials 94%
- A data-driven pipeline to extract potential side effects through co-prescription analysis: application to a cohort study of 2,010 patients taking hydroxychloroquine with an 11-year follow-up 93%
Similar papers in this journal
- Statistical inference for association studies using electronic health records: handling both selection bias and outcome misclassification 92%
- Improving Precision and Power in Randomized Trials for COVID-19 Treatments Using Covariate Adjustment, for Binary, Ordinal, and Time-to-Event Outcomes 92%
- Generalized Multi-SNP Mediation Intersection-Union Test 92%
Similar papers in this journal
- Incorporating data from multiple endpoints in the analysis of clinical trials: example from RSV vaccines 94%
- Assessing Direct and Spillover Effects of Intervention Packages in Network-Randomized Studies 93%
- Negative Control Exposures: Causal effect Identifiability and Use in Probabilistic-Bias and Bayesian Analyses with Unmeasured Confounders 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.