A Korean pangenome reference of 14 healthy individuals supports structural variant analysis in disease genomes
Shin, D.-H.; Jeon, J.; Joe, S.; Jeon, Y.; Yang, J. O.; Bhak, J.; Baek, S. A.; Byun, G.; Shin, E.-S.; Kwon, Y.; Choi, H.-J.; Kim, J.-H.; Haam, K.; Yoo, J.; Song, K. J.; Mok, J.; Jeon, S.; Jeong, H.; Bhak, J.
Show abstract
Here, we present the first graph-based Korean Pangenome Reference (K-PanRef), constructed from 14 healthy Korean individuals. K-PanRef comprises 13 high-quality diploid Korean genome assemblies (mean QV ~62.0) and KOREF1-G-TTAGGA, the first complete Korean reference genome. Integration of these assemblies generated a ~3.2-Gb pangenome graph containing ~39.3 million nodes and ~53.8 million edges, with the accumulation of common sequences (frequency [≥]10%) reaching a plateau. Additionally, K-PanRef contains ~4.3 million Korean-specific small variants and ~76.0 thousand Korean-specific SVs absent from the Chinese and human pangenome references, improving the representation of Korean genetic diversity relative to these references. To evaluate its utility for short-read-based SV analysis, we genotyped 75 whole-genome sequencing (WGS) samples, including 15 patients with early-onset myocardial infarction (MI). Although constructed entirely from healthy genomes, K-PanRef supported the identification of putative disease-relevant SVs in this exploratory application. K-PanRef-based genotyping identified ~95.6 thousand small variants and 820 SVs observed only in the early-onset MI samples. Among the early-onset MI-group SVs, 491 were absent from public databases, suggesting that they may represent previously unrecognized candidate variants related to early-onset MI. Of these, 164 SVs overlapped 134 genes, of which 89 had reported associations with 42 cardiovascular diseases or traits, including eight genes previously linked to MI. Together, these results establish K-PanRef as a valuable resource for representing Korean genetic diversity and enabling more comprehensive discovery of population-specific and novel putative disease-relevant variants from short-read sequencing data.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.