Back

10,239 whole genomes with multiomic and clinical health information as the Korean Multiomics Reference dataset

An, K.; Jeon, S.; Kwon, Y.; Choi, Y.; Yoon, C.; Jeon, Y.; Bhak, J.; Shin, D.-H.; Choi, H.-J.; Lee, H.; Kim, Y. J.; Shin, E.-S.; Ryu, H.; Blazyte, A.; Bolser, D.; Park, S.; Cho, J.; Joe, S.; Yang, J. O.; Jeon, J.; Kim, J.-H.; Kim, J.; Jung, D.; Cho, Y. S.; Chang, K.; Choo, E. H.; Kim, E.; Lee, S. Y.; Kim, W.; Kang, M. G.; Her, A.-Y.; Chon, S.; Woo, J.-T.; Rhee, S.; Lee, S.; Jin, H.-j.; Baek, Y.; Ban, H.-J.; Ahn, Y. M.; Rhee, S. J.; Kim, M. J.; Lee, S. Y.; Yang, C.-M.; Shim, S.-H.; Cho, S.-J.; Kim, S. G.; Jung, H.-T.; Ham, B.-J.; Choi, Y. Y.; Cheong, J.-H.; Kim, S.-K.; Phi, J. H.; Choi, S. A.; G

2026-01-26 genomics Community evaluation
10.1101/2025.11.17.688763 bioRxiv
Show abstract

We present Korea10K, the largest genomic dataset of the Korean population, comprising 10,239 high-coverage whole genomes (mean depth 30x) with matched multiomic profiles and phenotype data. Korea10K achieves complete and near-complete discovery of very rare and ultra-rare alleles, respectively, at 9,000 Korean genomes. This dataset provides the high-quality population-specific imputation panel, enabling accurate inference of low-frequency variants. Admixture analyses confirm the genetic homogeneity of the Korean population, despite its diverse Y-chromosomal, mitochondrial, and HLA repertoires. This pattern reflects a long and continuous lineage history characterized by persistent internal admixture and genomic homogenization over thousands of years on the Korean peninsula. We also identified 16.8 million genomic variants that directly modify CG sites by creating or abolishing CG dinucleotides, providing the population-scale evidence of coordinated genomic-epigenomic regulatory mechanism in Koreans.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.