Back

Integrating polygenic and transcriptional risk scores improves risk prediction of nine common diseases in the underrepresented Vietnamese population

Nguyen, S. V.; Pham, T. M.; Hoang, T. H.; Tran, T. T. H.; Vu, G. M.; Tran, M. H.; Nguyen, T. K.; Trinh, H. L.; Vu, H. T. T.; Pham, T. M.; Nghiem, D. T.; Pham, A. G.; Hoang, Y.; Giang, P. H.; Dao, D. X.; Luu, H. N.; Tran, T. H.; Nguyen, Q.; Truong, B.; Vo, N. S.

2025-12-09 genetic and genomic medicine
10.64898/2025.12.08.25341869 medRxiv
Show abstract

Polygenic risk scores (PRS) represent the cumulative impact of numerous common genomic variants, to predict clinical phenotypes and outcomes for individuals. However, PRS are typically derived from GWAS for populations of European origin, often resulting in reduced performance and their transferability to other underserved populations. In this study, we comprehensively analyzed 550 samples for nine common diseases in the Vietnamese population, including breast cancer (BC), colorectal cancer (CRC), gastric cancer (GC), chronic kidney disease (CKD), coronary artery disease (CAD), hyperlipidemia, osteoporosis, osteoarthritis, and Parkinsons disease (PD). Healthy control subjects were taken from the 1000 Vietnamese Genomes Project (VN1K). We evaluated seven advanced PRS algorithms using multiple GWAS datasets from both East Asian and European populations and identified the best performing method for each disease. PRS accuracy, assessed by incremental liability R-squared (incR2), ranged from 1.8% in CKD to 8.3% in CAD. The Area Under the Curve (AUC) ranged from 0.55 for CKD to 0.70 for CAD. Integrating with transcriptional risk scores (TRS), the PRS+TRS model led to a consistently increased incR2 across all nine diseases ranging from 1% to 15%. These findings offer valuable insights into the implementation of PRS+TRS for disease risk prediction in the Vietnamese population, where a similar approach would be applicable to other underrepresented populations.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.