Multi-ancestry gene-trait connection landscape using electronic health record (EHR) linked biobank data
Li, B.; Veturi, Y.; Lucas, A.; Bradford, Y.; Verma, S. S.; Verma, A.; Park, J.; Wei, W.-Q.; Feng, Q.; Namjou, B.; Kiryluk, K.; Kullo, I.; Luo, Y.; Pividori, M.; Im, H. K.; Greene, C. S.; Ritchie, M. D.
Show abstract
Understanding genetic factors of complex traits across ancestry groups holds a key to improve the overall health care quality for diverse populations in the United States. In recent years, multiple electronic health record-linked (EHR-linked) biobanks have recruited participants of diverse ancestry backgrounds; these biobanks make it possible to obtain phenome-wide association study (PheWAS) summary statistics on a genome-wide scale for different ancestry groups. Moreover, advancement in bioinformatics methods provide novel means to accelerate the translation of basic discoveries to clinical utility by integrating GWAS summary statistics and expression quantitative trait locus (eQTL) data to identify complex trait-related genes, such as transcriptome-wide association study (TWAS) and colocalization analyses. Here, we combined the advantages of multi-ancestry biobanks and data integrative approaches to investigate the multi-ancestry, gene-disease connection landscape. We first performed a phenome-wide TWAS on Electronic Medical Records and Genomics (eMERGE) III network participants of European ancestry (N = 68,813) and participants of African ancestry (N = 12,658) populations, separately. For each ancestry group, the phenome-wide TWAS tested gene-disease associations between 22,535 genes and 309 curated disease phenotypes in 49 primary human tissues, as well as cross-tissue associations. Next, we identified gene-disease associations that were shared across the two ancestry groups by combining the ancestry-specific results via meta-analyses. We further applied a Bayesian colocalization method, fastENLOC, to prioritize likely functional gene-disease associations with supportive colocalized eQTL and GWAS signals. We replicated the phenome-wide gene-disease analysis in the analogous Penn Medicine BioBank (PMBB) cohorts and sought additional validations in the PhenomeXcan UK Biobank (UKBB) database, PheWAS catalog, and systematic literature review. Phenome-wide TWAS identified many proof-of-concept gene-disease associations, e.g. FTO-obesity association (p = 7.29e-15), and numerous novel disease-associated genes, e.g. association between GATA6-AS1 with pulmonary heart disease (p = 4.60e-10). In short, the multi-ancestry, gene-disease connection landscape provides rich resources for future multi-ancestry complex disease research. We also highlight the importance of expanding the size of non-European ancestry datasets and the potential of exploring ancestry-specific genetic analyses as these will be critical to improve our understanding of the genetic architecture of complex disease.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Leveraging genomic diversity for discovery in an EHR-linked biobank: the UCLA ATLAS Community Health Initiative 96%
- Defining and Reducing Variant Classification Disparities 94%
- An atlas connecting shared genetic architecture of human diseases and molecular phenotypes provides insight into COVID-19 susceptibility 94%
Similar papers in this journal
- Cancer PRSweb - an Online Repository with Polygenic Risk Scores (PRS) for Major Cancer Traits and Their Phenome-wide Exploration in Two Independent Biobanks 96%
- Assessing digital phenotyping to enhance genetic studies of human diseases 95%
- The Phenotype-Genotype Reference Map: Improving biobank data science through replication. 95%
Similar papers in this journal
- Unsupervised representation learning improves genomic discovery and risk prediction for respiratory and circulatory functions and diseases 95%
- Genome-wide analysis in 756,646 individuals provides first genetic evidence that ACE2 expression influences COVID-19 risk and yields genetic risk scores predictive of severe disease 94%
- Large scale genome-wide association study in a Japanese population identified 45 novel susceptibility loci for 22 diseases 94%
Similar papers in this journal
- The genetic and phenotypic correlates of mtDNA copy number in a multi-ancestry cohort 95%
- Multivariate adaptive shrinkage improves cross-population transcriptome prediction for transcriptome-wide association studies in underrepresented populations 95%
- Disease-specific prioritization of non-coding GWAS variants based on chromatin accessibility 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.