Back

Constructing genotype and phenotype network helps reveal disease heritability and phenome-wide association studies

Cao, X.; Zhu, L.; Liang, X.; Zhang, S.; Sha, Q.

2023-11-20 genetic and genomic medicine
10.1101/2023.11.14.23297400 medRxiv
Show abstract

Analyses of a bipartite Genotype and Phenotype Network (GPN), linking the genetic variants and phenotypes based on statistical associations, provide an integrative approach to elucidate the complexities of genetic relationships across diseases and identify pleiotropic loci. In this study, we first assess contributions to constructing a well-defined GPN with a clear representation of genetic associations by comparing the network properties with a random network, including connectivity, centrality, and community structure. Next, we construct network topology annotations of genetic variants that quantify the possibility of pleiotropy and apply stratified linkage disequilibrium (LD) score regression to 12 highly genetically correlated phenotypes to identify enriched annotations. The constructed network topology annotations are informative for disease heritability after conditioning on a broad set of functional annotations from the baseline-LD model. Finally, we extend our discussion to include an application of bipartite GPN in phenome-wide association studies (PheWAS). The community detection method can be used to obtain a priori grouping of phenotypes detected from GPN based on the shared genetic architecture, then jointly test the association between multiple phenotypes in each network module and one genetic variant to discover the cross-phenotype associations and pleiotropy. Significance thresholds for PheWAS are adjusted for multiple testing by applying the false discovery rate (FDR) control approach. Extensive simulation studies and analyses of 633 electronic health record (EHR)-derived phenotypes in the UK Biobank GWAS summary dataset reveal that most multiple phenotype association tests based on GPN can well-control FDR and identify more significant genetic variants compared with the tests based on UK Biobank categories.

Published in BMC Genomics · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 16%
11.7%
2
Nature Computational Science
55 papers in training set
Top 0.1%
11.7%
3
Bioinformatics
1204 papers in training set
Top 3%
9.6%
4
The American Journal of Human Genetics
234 papers in training set
Top 0.6%
7.7%
5
Scientific Reports
3612 papers in training set
Top 9%
7.1%
6
Human Genetics and Genomics Advances
84 papers in training set
Top 0.4%
4.8%
50% of probability mass above
7
Cell Genomics
172 papers in training set
Top 0.8%
4.3%
8
Patterns
78 papers in training set
Top 0.4%
4.0%
9
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.4%
10
npj Genomic Medicine
36 papers in training set
Top 0.1%
3.4%
11
Nature Genetics
286 papers in training set
Top 2%
2.7%
12
iScience
1154 papers in training set
Top 10%
2.6%
13
Frontiers in Genetics
230 papers in training set
Top 2%
2.4%
14
Communications Biology
993 papers in training set
Top 11%
2.1%
15
PLOS Genetics
862 papers in training set
Top 6%
1.9%
16
Genome Biology
637 papers in training set
Top 6%
1.7%
17
Genome Medicine
183 papers in training set
Top 4%
1.1%
18
Genetic Epidemiology
55 papers in training set
Top 0.5%
1.1%
19
PLOS Computational Biology
1863 papers in training set
Top 17%
1.1%
20
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 0.8%
1.1%
21
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
1.1%
22
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 36%
1.1%
23
The Innovation
13 papers in training set
Top 0.2%
1.0%
24
European Journal of Human Genetics
58 papers in training set
Top 1%
0.8%
25
eLife
5828 papers in training set
Top 70%
0.6%
26
Nature Human Behaviour
95 papers in training set
Top 3%
0.6%
27
npj Digital Medicine
118 papers in training set
Top 4%
0.6%