Back

An Expectation and Maximization Algorithm for Multivariate Genome-wide Association Studies (EMmvGWAS)

Teng, C.-S.; Wang, X.; Liu, C.; Wang, Q.; Cui, Y.; Xu, S.

2025-12-05 genomics
10.64898/2025.12.02.691906 bioRxiv
Show abstract

Genome-wide association studies (GWAS) commonly focus on one quantitative trait at a time, even when multiple traits are collected. However, joint analysis of multiple traits can increase the power to detect genetic associations by leveraging trait correlations. Multivariate GWAS is particularly important for identifying pleiotropic effects and uncovering shared genetic architecture underlying complex traits. Despite its great potential, multivariate GWAS methods face substantial computational challenges due to the high dimensionality of polygenic covariance structures and the cost of scanning genome-wide markers. We present EMmvGWAS, an efficient computational framework implemented as an R, which dramatically reduces computation time for multivariate GWAS involving a moderate number of traits. The algorithm estimates the genetic and environmental covariance matrices from a pure polygenic model using an Expectation-Maximization (EM) algorithm. To improve computational efficiency during genome-wide scanning, we use a semi-exact method that fixes the estimated covariance matrices and treats their ratio, defined as the genetic covariance matrix multiplied by the inverse of the environmental covariance matrix, as a known constant across all markers. This approach enables closed-form solutions for marker effects at each locus, dramatically reducing the computation time without compromising statistical power. We demonstrate the scalability and performance of our method through both simulation studies and real dataset analyses from rice, mouse and human populations. The method is implemented in R and is designed to support multivariate GWAS involving a moderate number of traits, while also accommodating univariate analyses as a special case. The R package is available at https://github.com/Jason-Teng/EMmvGWAS.

Published in GENETICS (predicted rank #5) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.