Privacy-Enhancing Sequential Learning under Heterogeneous Selection Bias in Multi-Site EHR Data
Kundu, R.; Shi, X.; Patel, K. K.; Ohno-Machado, L.; Salvatore, M.; Song, P. X. K.; Mukherjee, B.
Show abstract
ObjectiveTo develop privacy-enhancing statistical methods for estimation of binary disease risk model association parameters across multiple electronic health record (EHR) sites with heterogeneous selection mechanisms, without sharing raw individual-level data. We illustrate their utility through a cross-biobank analysis of smoking and 97 cancer subtypes using data from the NIH All of Us (AOU) and the Michigan Genomics Initiative (MGI). Materials and MethodsLarge-scale biobanks often follow heterogeneous recruitment strategies and store data in separate cloud-based platforms, making centralized algorithms infeasible. To address this, we propose two decentralized sequential estimators namely, Sequential Pseudo-likelihood (SPL) and Sequential Augmented Inverse Probability Weighting (SAIPW) that leverage external population-level information to adjust for selection bias, with valid variance estimation. SAIPW additionally protects against misspecification of the selection model using flexible machine learning based auxiliary outcome models. We compare SPL and SAIPW with the existing Sequential Unweighted (SUW) estimator and with centralized and meta learning extensions of IPW and AIPW in simulations under both correctly specified and misspecified selection mechanisms. We apply the methods to harmonized data from MGI (n = 50,935) and AOU (n = 241,563) to estimate smoking-cancer associations. ResultsIn simulations, SUW exhibited substantial bias and poor coverage. SPL and SAIPW yielded unbiased estimates with valid coverage probabilities under correct model specification, with SAIPW remaining robust under selection model misspecification. Both approaches showed no notable efficiency loss relative to centralized methods. Meta-learning methods were efficient for large sites but failed in settings with small cohort sizes and rare outcome prevalence. In real-data analysis, strong associations were consistently identified between smoking and cancers of the lung, bladder, and larynx, aligning with established epidemiological evidence. ConclusionOur framework enables valid, privacy-enhancing inference across EHR cohorts with heterogeneous selection, supporting scalable, decentralized research using real-world data.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Double Machine Learning Approach for the Evaluation of COVID-19 Vaccine Effectiveness under the Test-Negative Design: Analysis of Québec Administrative Data 96%
- Bias reduction and inference for electronic health record data under selection and phenotype misclassification: three case studies 96%
- Probabilistic Cause-of-disease Assignment using Case-control Diagnostic Tests: A Latent Variable Regression Approach 96%
Similar papers in this journal
- Generalized Multi-SNP Mediation Intersection-Union Test 95%
- Operating Characteristics of the Rank-Based Inverse Normal Transformation for Quantitative Trait Analysis in Genome-Wide Association Studies 95%
- A Novel Penalized Inverse-Variance Weighted Estimator for Mendelian Randomization with Applications to COVID-19 Outcomes 94%
Similar papers in this journal
- A Two-Sample Robust Bayesian Mendelian Randomization Method Accounting for Linkage Disequilibrium and Idiosyncratic Pleiotropy with Applications to the COVID-19 Outcome 95%
- Statistics to prioritize rare variants in family-based sequencing studies with disease subtypes 95%
- Identifying causal genotype-phenotype relationships for population-sampled parent-child trios 94%
Similar papers in this journal
- A mixed-model approach for powerful testing of genetic associations with cancer risk incorporating tumor characteristics 95%
- Fast Lasso method for Large-scale and Ultrahigh-dimensional Cox Model with applications to UK Biobank 95%
- Survival Analysis on Rare Events Using Group-Regularized Multi-Response Cox Regression 95%
Similar papers in this journal
- Joint Modeling of Longitudinal Biomarker and Survival Outcomes with the Presence of Competing Risk in Nested Case-Control Studies with Application to the TEDDY Microbiome Dataset 95%
- Subset scanning for multi-trait analysis using GWAS summary statistics 95%
- High-dimensional Biomarker Identification for Scalable and Interpretable Disease Prediction via Machine Learning Models 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.