An Expectation and Maximization Algorithm for Multivariate Genome-wide Association Studies (EMmvGWAS)
Teng, C.-S.; Wang, X.; Liu, C.; Wang, Q.; Cui, Y.; Xu, S.
Show abstract
Genome-wide association studies (GWAS) commonly focus on one quantitative trait at a time, even when multiple traits are collected. However, joint analysis of multiple traits can increase the power to detect genetic associations by leveraging trait correlations. Multivariate GWAS is particularly important for identifying pleiotropic effects and uncovering shared genetic architecture underlying complex traits. Despite its great potential, multivariate GWAS methods face substantial computational challenges due to the high dimensionality of polygenic covariance structures and the cost of scanning genome-wide markers. We present EMmvGWAS, an efficient computational framework implemented as an R, which dramatically reduces computation time for multivariate GWAS involving a moderate number of traits. The algorithm estimates the genetic and environmental covariance matrices from a pure polygenic model using an Expectation-Maximization (EM) algorithm. To improve computational efficiency during genome-wide scanning, we use a semi-exact method that fixes the estimated covariance matrices and treats their ratio, defined as the genetic covariance matrix multiplied by the inverse of the environmental covariance matrix, as a known constant across all markers. This approach enables closed-form solutions for marker effects at each locus, dramatically reducing the computation time without compromising statistical power. We demonstrate the scalability and performance of our method through both simulation studies and real dataset analyses from rice, mouse and human populations. The method is implemented in R and is designed to support multivariate GWAS involving a moderate number of traits, while also accommodating univariate analyses as a special case. The R package is available at https://github.com/Jason-Teng/EMmvGWAS.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Interpretable Artificial Neural Networks incorporating Bayesian Alphabet Models for Genome-wide Prediction and Association Studies 95%
- A Multiple-trait Bayesian Variable Selection Regression Method for Integrating Phenotypic Causal Networks in Genome-Wide Association Studies 95%
- gJLS2: A generalized joint location and scale analysis tool for X-inclusive genome-wide discoveries 95%
Similar papers in this journal
- Joint Modeling of Effect Sizes for Two Correlated Traits: Characterizing Trait Properties to Enhance Polygenic Risk Prediction 96%
- ADELLE: A global testing method for Trans-eQTL mapping 96%
- A Fast and Scalable Framework for Large-scale and Ultrahigh-dimensional Sparse Regression with Application to the UK Biobank 96%
Similar papers in this journal
- Bayesian Hierarchical Hypothesis Testing in Large-Scale Genome-Wide Association Analysis 97%
- Using encrypted genotypes and phenotypes for collaborative genomic analyses to maintain data confidentiality 96%
- Estimating SNP heritability in presence of population substructure in large biobank-scale data 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.