A Bayesian network-based framework to uncover the causal effects of genes on complex traits based on GWAS data
YIN, L.; FENG, Y.; LAU, A.; QIU, J.; Sham, P.-C.; SO, H.-C.
Show abstract
Deciphering the relationships between genes and complex traits could help us better understand the biological mechanisms leading to phenotypic variations and disease onset. Univariate gene-based analyses are widely used to characterize gene-phenotype relationships, but are subject to the influence of confounders. Furthermore, while some genes directly contribute to traits variations, others may exert their effects through other genes. How to quantify individual genes direct and indirect effects on complex traits remains an important yet challenging question. We presented a novel framework to decipher the total and direct causal effects of individual genes using imputed gene expression data from GWAS and raw gene expression from GTEx. The study was partially motivated by the quest to differentiate "core" genes (genes with direct causal effect on the phenotype) from "peripheral" ones. Our proposed framework is based on a Bayesian network (BN) approach, which produces a directed graph showing the relationship between genes and the phenotype. The approach aims to uncover the overall causal structure, to examine the role of individual genes and quantify the direct and indirect effects by each gene. An important advantage and novelty of the proposed framework is that it allows gene expression and disease trait(s) to be evaluated in different samples, significantly improving the flexibility and applicability of the approach. It uses IDA and jointIDA incorporating a novel p-value-based regularization approach to quantify the causal effects (including total causal effects, direct causal effects, and medication effects) of genes. The proposed approach can be extended to decipher the joint causal network of 2 or more traits, and has high specificity and precision (a.k.a., positive predictive value), making it particularly useful for selecting genes for follow-up studies. We verified the feasibility and validity of the proposed framework by extensive simulations and applications to 52 traits in the UK Biobank (UKBB). Split-half replication and stability selection analyses were performed to demonstrate the accuracy and efficiency of our proposed method to identify causally relevant genes. The identified (direct) causal genes were found to be significantly enriched for genes highlighted in the OpenTargets database, and the enrichment was stronger than achieved by conventional univariate gene-based tests. Encouragingly, many enriched pathways were supported by the literature, and some of the enriched drugs have been tested or used to treat patients in clinical practice. Our proposed framework provides powerful a way to prioritize genes with large direct or indirect causal effects and to quantify the importance of such genes.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BayesKAT: Bayesian Optimal Kernel-based Test for genetic association studies reveals joint genetic effects in complex diseases 95%
- CoRegNet: Unraveling Gene Co-regulation Networks from Public RNA-Seq Repositories Using a Beta-Binomial Statistical Model 95%
- Enhancing single-cell cellular state inference by incorporating molecular network features 95%
Similar papers in this journal
- netMUG: a novel network-guided multi-view clustering workflow for dissecting genetic and facial heterogeneity 94%
- GMQN: A reference-based method for correcting batch effects as well as probes bias in HumanMethylation BeadChip 92%
- Tensor decomposition-Based Unsupervised Feature Extraction Applied to Single-Cell Gene Expression Analysis 92%
Similar papers in this journal
- Explainable deep transfer learning model for disease risk prediction using high-dimensional genomic data 96%
- A Generalized Higher-order Correlation Analysis Framework for Multi-Omics Network Inference 96%
- DNFE: Directed-network flow entropy for detecting the tipping points during biological processes 95%
Similar papers in this journal
- A novel method for multiple phenotype association studies based on genotype and phenotype network 96%
- Network Assisted Analysis of De Novo Variants Using Protein-Protein Interaction Information Identified 46 Candidate Genes for Congenital Heart Disease 95%
- Searching across-cohort relatives in 54,092 GWAS samples via encrypted genotype regression 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.