Back

Large-scale imputation models for multi-ancestry proteome-wide association analysis

Wu, C.; Zhang, Z.; Yang, X.; Zhao, B.

2023-10-09 genetics
10.1101/2023.10.05.561120 bioRxiv
Show abstract

Proteome-wide association studies (PWAS) decode the intricate proteomic landscape of biological mechanisms for complex diseases. Traditional PWAS model training relies heavily on individual-level reference proteomes, restricting its capacity to harness the emerging summary-level protein quantitative trait loci (pQTL) data in the public domain. Here we introduced BLISS, a novel framework to train protein imputation models using only pQTL summary statistics. By leveraging extensive pQTL data from the UK Biobank, deCODE, and ARIC studies, we applied BLISS to develop large-scale European PWAS models covering 5,779 unique proteins. We further extended BLISS to integrate with small-scale non-European individual-level datasets, enabling the development of models tailored to Asian and African ancestries. We validated the performance of BLISS models through a systematic multi-ancestry analysis of over 2,500 phenotypes across five major genetic data resources. The newly identified protein-phenotype associations offer valuable insights into their cross-ancestry transferability, the contributions of different proteomic platforms, and the complementary perspectives provided by distinct genomic mapping approaches relevant to drug discovery. The developed models and data resources are freely available at https://www.gcbhub.org/.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.