Back

A Federated and Privacy-Preserving Framework for Large-Scale Genome-Wide Association Studies with Mixed-Effects Models

Suo, X.; Xue, F.; Zhao, Y.

2025-12-19 genomics
10.64898/2025.12.16.693409 bioRxiv
Show abstract

Genome-wide association studies (GWAS) increasingly rely on large-scale data integration to achieve the statistical power necessary to detect variants with weak effects. However, genomic data are typically siloed across institutions, and privacy constraints often preclude centralized analysis. While federated learning (FL) offers a viable alternative by enabling cross-site computation without sharing individual-level data, applying mixed models, which are essential for correcting population structure, in a distributed setting remains a challenge in terms of statistical accuracy and computational scalability. Here, we present a federated mixed-model framework for GWAS that achieves high fidelity to centralized analyses while maintaining efficiency at biobank scale. Building on mixed-model theory and distributed optimization, we introduce algorithms for continuous (FedLMM) and binary (FedGLMM) traits that perform parameter estimation and association testing through site-local computation and aggregation of intermediate statistics. Comprehensive simulations spanning varied sample sizes and genomic densities demonstrate that our methods closely mirror centralized benchmarks (fastGWA and fastGWA-GLMM). Effect-size estimates exhibit near-perfect correlation, and over 99% of significant loci are recovered with well-controlled type I error rates. Empirical analyses on [~]100,000 UK Biobank participants further confirm that the framework delivers consistent inference while sustaining high computational performance. This work establishes a practical, open-source, and statistically reliable federated solution for large-scale GWAS, resolving the tension between data privacy and the need for statistical power in modern genomics.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.