Convex approaches to isolate the shared and distinct genetic structures of subphenotypes in heterogeneous complex traits
Banerjee, S.; O'Connell, S.; Colbert, S.; Mullins, N.; Knowles, D. A.
Show abstract
Groups of complex diseases, such as coronary heart diseases, neuropsychiatric disorders, and cancers, often display overlapping clinical symptoms and pharmacological treatments. The shared associations of genetic variants across diseases has the potential to explain their underlying biological processes, but this remains poorly understood. To address this, we model the matrix of summary statistics of trait-associated genetic variants as the sum of a low-rank component - representing shared biological processes - and a sparse component, representing unique processes and arbitrarily corrupted or contaminated components. We introduce Clorinn, an open-source Python software that uses convex optimization algorithms to recover these components by minimizing a weighted combination of the nuclear norm and of the L1 norm. Among others, Clorinn provides two significant benefits: (a) Convex optimization guarantees reproducibility of the components, and (b) The low-rank "uncor-rupted" matrix allows robust singular value decomposition (SVD) and principal component analysis (PCA), which are otherwise highly sensitive to outliers and noise in the input matrix. In extensive simulations, we observe that Clorinn outperforms state-of-the-art approaches in capturing the shared latent factors across phenotypes. We apply Clorinn to estimate 200 latent factors from GWAS summary data of 2,110 phenotypes measured in European-ancestry Pan-UK BioBank individuals (N = 420,531) and 14 psychiatric disorders.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Fast and Accurate Bayesian Polygenic Risk Modeling with Variational Inference 96%
- Sparse modeling of interactions enables fast detection of genome-wide epistasis in biobank-scale studies 96%
- Welch-weighted Egger regression reduces false positives due to correlated pleiotropy in Mendelian randomization 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Co-expression-wide association studies link genetically regulated interactions with complex traits 96%
- Fast Kernel-based Association Testing of non-linear genetic effects for Biobank-scale data 96%
- Simultaneous estimation of bi-directional causal effects and heritable confounding from GWAS summary statistics 96%
Similar papers in this journal
- Mendelian randomization accounting for correlated and uncorrelated pleiotropic effects using genome-wide summary statistics. 95%
- Leveraging a machine learning derived surrogate phenotype to improve power for genome-wide association studies of partially missing phenotypes in population biobanks 94%
- LDAK-KVIK performs fast and powerful mixed-model association analysis of quantitative and binary phenotypes 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.