Back

A multivariate approach to joint testing of main genetic and gene-environment interaction effects

Mishra, S.; Majumdar, A.

2024-05-08 genomics
10.1101/2024.05.06.592645 bioRxiv
Show abstract

Gene-environment (GxE) interactions crucially contribute to complex phenotypes. The statistical power of a GxE interaction study is limited mainly due to weak GxE interaction effect sizes. To utilize the individually weak GxE effects to improve the discovery of associated genetic loci, Kraft et al. [1] proposed a joint test of the main genetic and GxE effects for a univariate phenotype. We develop a testing procedure to evaluate combined genetic and GxE effects on a multivariate phenotype to enhance the power by merging pleiotropy in the main genetic and GxE effects. We base the approach on a general linear hypothesis testing framework for a multivariate regression for continuous phenotypes. We implement the generalized estimating equations (GEE) technique under the seemingly unrelated regressions (SUR) setup for binary or mixed phenotypes. We use extensive simulations to show that the test for joint multivariate genetic and GxE effects outperforms the univariate joint test of genetic and GxE effects and the test for multivariate GxE effect concerning power when there is pleiotropy. The test produces a higher power than the test for multivariate main genetic effect for a weak genetic and substantial GxE effect. For more prominent genetic effects, the latter performs better with a limited increase in power. Overall, the multivariate joint approach offers high power across diverse simulation scenarios. We apply the methods to lipid phenotypes with sleep duration as an environmental factor in the UK Biobank. The proposed approach identified six independent associated genetic loci missed by other competing methods.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.