Back

An Improved Linear Mixed Model for Multivariate Genome-Wide Association Studies

Wang, D.; Teng, J.; Zhao, C.; Li, W.; Tang, H.; Fan, X.; Zhang, Q.; Ning, C.

2022-02-22 bioinformatics
10.1101/2022.02.21.481252 bioRxiv
Show abstract

Current methods of multivariate analysis require complete multivariate phenotypes from each individual and have a computational time complexity of O(n2) per SNP, where n is the sample size. We develop an efficient genomic multivariate analysis tool (GMAT) for genome-wide association studies of multiple correlated traits. The new method can handle incomplete multivariate data with missing records and reduce the time complexity to O(n) per SNP. Simulation studies based on known genotypes and phenotypes of actual populations show that GMAT has increased the statistical power with a proper control of false positivity for association studies compared to the conventional linear mixed model (LMM) that removes individuals with incomplete records. Applications to a balanced donkey data and an unbalanced yeast data show that the computational efficiency of the new method has been increased about tens of times faster than the conventional LMM analysis. The GMAT package can be downloaded at https://github.com/chaoning/GMAT.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.