mixWAS: An efficient distributed algorithm for mixed-outcomes genome-wide association studies
Li, R.; Benz, L.; Duan, R.; Denny, J.; Hakonarson, H.; Mosley, J.; Smoller, J. W.; Wei, W.-Q.; Ritchie, M. D.; Moore, J. H.; Chen, Y.
Show abstract
In cross-cohort studies, integrating diverse datasets, such as electronic health records (EHRs), is both essential and challenging due to cohort-specific variations, distributed data storage, and data privacy concerns. Traditional methods often require data pooling or complex data harmonization, which can reduce efficiency and limit the scope of cross-cohort learning. We introduce mixWAS, a one-shot, lossless algorithm that efficiently integrates distributed EHR datasets via summary statistics. Unlike existing approaches, mixWAS preserves cohort-specific covariate associations and supports simultaneous mixed-outcome analyses. Simulations demonstrate that mixWAS outperforms conventional methods in accuracy and efficiency across various scenarios. Applied to EHR data from seven cohorts in the US, mixWAS identified 4,530 significant cross-cohort genetic associations among traits such as blood lipids, BMI, and circulatory diseases. Validation with an independent UK EHR dataset confirmed 97.7% of these associations, underscoring the algorithms robustness. By enabling lossless cross-cohort integration, mixWAS improves the precision of multi-outcome analyses and expands the potential for actionable insights in healthcare research. The bigger pictureCross-cohort integration of electronic health record (EHR) datasets is critical for advancing genomic discovery but remains hindered by privacy concerns, cohort heterogeneity, and computational limitations. Traditional meta-analysis and federated methods either lose power or cannot fully model multiple mixed-outcome traits across distributed datasets. To address this, we developed mixWAS, a one-shot, lossless algorithm for integrating summary statistics across cohorts without sharing individual-level data. mixWAS simultaneously models binary and continuous outcomes, accounts for site-specific covariate heterogeneity, and requires only a single communication step between sites. Through extensive simulations and real data analyses, mixWAS consistently outperformed traditional Phenome-Wide Association Studies (PheWAS) and other multi-trait approaches in detecting multi-phenotype associations (MPAs). eyond genetic applications, mixWAS offers a general framework for distributed analysis of mixed-outcome data, with broad potential across biomedicine, public health, and other fields requiring privacy- preserving data integration. HighlightsO_LImixWAS enables lossless, one-shot cross-cohort integration of summary statistics C_LIO_LISimultaneously models binary and continuous outcomes across distributed datasets C_LIO_LIOutperforms PheWAS in detecting multi-phenotype associations (MPA) C_LIO_LIOffers a general framework for distributed analysis of mixed-outcome data, C_LI
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identification of putative causal loci in whole-genome sequencing data via knockoff statistics 96%
- OTTERS: A powerful TWAS framework leveraging summary-level reference data 96%
- Testing and controlling for horizontal pleiotropy with the probabilistic Mendelian randomization in transcriptome-wide association studies 96%
Similar papers in this journal
- Summary statistics from large-scale gene-environment interaction studies for re-analysis and meta-analysis 96%
- Incorporating family disease history and controlling case-control imbalance for population based genetic association studies 96%
- Uncovering genetic associations in the human diseasome using an endophenotype-augmented disease network 95%
Similar papers in this journal
- Deep transfer learning provides a Pareto improvement for multi-ancestral clinico-genomic prediction of diseases 94%
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 94%
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 94%
Similar papers in this journal
- Leveraging TOPMed Imputation Server and Constructing a Cohort-Specific Imputation Reference Panel to Enhance Genotype Imputation among Cystic Fibrosis Patients 95%
- Multivariate adaptive shrinkage improves cross-population transcriptome prediction for transcriptome-wide association studies in underrepresented populations 95%
- Polygenic risk score prediction accuracy convergence 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.