Back

Genomic atlas of 7,000 plasma proteins and their associations with diseases and traits in East Asian populations

Pozarickij, A.; Wang, B.; Mohamed, A.; Lin, K.; Morris, S.; Kartsonaki, C.; Wright, N.; Fry, H.; Chen, Y.; Du, H.; Bennett, D.; Yang, L.; Avery, D.; Schmidt, D. V.; Lin, L.; Lv, J.; Yu, C.; Sun, D.; Pei, P.; Chen, J.; Hill, M.; Peto, R.; Collins, R.; Clarke, R.; Millwood, I. Y.; Chen, Z.; Walters, R. G.

2026-02-06 epidemiology
10.64898/2026.02.05.26345625 medRxiv
Show abstract

Proteogenomic studies integrating genetic, molecular, and phenotypic data have transformed target discovery, yet remain heavily biased toward European populations. Here, we present a large-scale proteogenomic atlas in a non-European population, analysing 7,289 plasma proteins profiled by SomaScan v4.1 in 3,965 Chinese adults. Genome-wide association analyses identified 3,212 protein quantitative trait loci (pQTLs), including 1,092 proteins with a cis-pQTL. Integrating these data with East Asian phenotypes and disease outcomes, we performed proteome-wide phenome scans and identified 7,936 protein-phenotype associations with strong colocalization support (PP.H4 > 0.8). Mendelian randomisation analyses using cis-pQTL instruments further prioritised 1,975 protein-phenotype associations, with 645 high-confidence pairs supported by both colocalisation and causal inference. Notably, we identified ancestry-specific pQTLs that contributed to associations undetectable in European studies alone. These associations organised into coherent biological networks, most prominently involving lipid metabolism and cardiovascular disease. Together, this study expands the global proteogenomic landscape and establishes a publicly valuable atlas of genetically anchored protein-phenotype relationships, providing a foundational resource for future genetic, functional, and translational studies, including drug-target prioritisation and risk-benefit assessment.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.