Back

European-derived coronary artery disease polygenic scores over-flag genetic risk in Vietnamese and Southeast Asian populations: a multi-score analysis in 1000 Genomes

Hoang, Q. P.; Le, T. X.; Doan, D. D.

2026-07-15 genetic and genomic medicine
10.64898/2026.07.10.26357796 medRxiv
Show abstract

Background. Polygenic scores (PRS) for coronary artery disease (CAD) are derived almost entirely from European-ancestry data. Their portability to Southeast Asian populations, including the Vietnamese, is largely uncharacterised and clinically consequential when scores are used with risk thresholds. Methods. We evaluated four independent European-derived CAD scores from the PGS Catalog (PGS000058, PGS000349, PGS002809, PGS004198; 70 - 5,723 variants) in 2,504 individuals from the 1000 Genomes Project, focusing on the Vietnamese Kinh (KHV) and Dai (CDX) samples. Per-individual scores were computed with PLINK2 and standardised. We assessed (i) the cross-ancestry distribution (calibration) and (ii) a clinically-relevant consequence: the proportion of each population flagged high genetic risk when the European top-20% threshold is applied (20% if perfectly calibrated). Results. For the primary score (PGS000058) the standardised PRS differed across super-populations (ANOVA F(4, 2499) = 121.1, p < 0.001); the Vietnamese Kinh mean was +0.47 SD above the European mean (Welch t = 7.77, p = 2.0 x 10^ -14). Applying the European top-20% high-risk threshold, the fraction of Vietnamese Kinh flagged ranged from 22.2% to 57.6% across the four scores, and of Dai from 21.5% to 43.0%, versus the intended 20%. Three of the four scores over-flagged Vietnamese (25-58%); the largest score (PGS004198) was approximately calibrated for East/Southeast Asians ([~]22%) but markedly over-flagged Africans (69.3%). Conclusions. European-derived CAD polygenic scores are inconsistently calibrated in Vietnamese and other Southeast Asian samples, and most substantially over-flag high genetic risk when a European threshold is applied. The magnitude and even the direction of miscalibration depend on the specific score, so no such score can be assumed transferable without local validation and recalibration. Distribution shift bounds, but does not by itself quantify, loss of predictive accuracy, which requires phenotyped data.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Circulation: Genomic and Precision Medicine
48 papers in training set
Top 0.1%
12.8%
2
Human Genetics and Genomics Advances
84 papers in training set
Top 0.1%
12.8%
3
European Journal of Human Genetics
58 papers in training set
Top 0.1%
9.9%
4
The American Journal of Human Genetics
234 papers in training set
Top 0.5%
9.0%
5
Genome Medicine
183 papers in training set
Top 0.5%
5.5%
50% of probability mass above
6
Scientific Reports
3612 papers in training set
Top 23%
4.4%
7
Nature Communications
5641 papers in training set
Top 35%
3.3%
8
Genetic Epidemiology
55 papers in training set
Top 0.2%
3.3%
9
Genes
144 papers in training set
Top 0.9%
2.8%
10
Human Genomics
21 papers in training set
Top 0.1%
2.8%
11
PLOS ONE
5266 papers in training set
Top 41%
2.7%
12
Frontiers in Genetics
230 papers in training set
Top 2%
2.5%
13
Human Genetics
28 papers in training set
Top 0.2%
2.4%
14
PLOS Genetics
862 papers in training set
Top 6%
2.1%
15
Human Molecular Genetics
141 papers in training set
Top 2%
1.5%
16
Genetics in Medicine
78 papers in training set
Top 0.8%
1.1%
17
GENETICS
483 papers in training set
Top 3%
1.1%
18
eLife
5828 papers in training set
Top 57%
1.1%
19
Communications Biology
993 papers in training set
Top 21%
1.1%
20
BMC Medicine
176 papers in training set
Top 5%
0.9%
21
Atherosclerosis
30 papers in training set
Top 0.7%
0.9%
22
Communications Medicine
113 papers in training set
Top 5%
0.9%
23
JAMA
18 papers in training set
Top 0.5%
0.6%
24
International Journal of Epidemiology
88 papers in training set
Top 2%
0.6%