A Complete Telomere-to-Telomere Diploid Reference Genome for Indian Population
Sarashetti, P.; Lipovac, J.; Jia, Q.; Wang, L.; Li, Z.; Vrcek, L.; Yao, F.; Sia, Y.; Lin, D.; Zhang, X.; Muliaditan, D.; Low, H. M.; Leong, S. T.; Khor, C. C.; Wong, E.; Zheng, W.; Krizanovic, K.; Chambers, J. C.; Hwang, W. Y. K.; Tan, P.; Sikic, M.; Liu, J.
Show abstract
Human reference genomes have been instrumental in advancing genomic and biomedical research, but South and Southeast Asian populations are underrepresented, despite accounting for a large proportion of world population. As a part of effort on generating reference genomes for these populations, we present the first gapless, telomere-to-telomere (T2T) diploid genome assembly created by using a trio sample set of Indian ancestry (I002C), with NG50 of 154.89 Mb and 146.27 Mb for the maternal and paternal haplotypes, including the fully assembled rDNA array for the maternal chromosome 21 and Y chromosome. With the Merqury QVs of 82.05, 83.08 and 82.64 for the maternal, paternal and haploid assemblies respectively, I002C represents the highest-quality human genome assembled in both diploid and haploid forms to date. Compared to CHM13, the I002C genome displays substantial sequence diversity, resulting in 14,943 structural variants, including 3,236 novel variants absent from public databases. Analysis of trio-phased haplotypes further revealed elevated inter-haplotype divergence within centromeric and subtelomeric regions, along with identification of differentially methylated regions (DMRs) as candidates for novel imprinting loci. As a result of substantial SVs between them, I002C is a more suitable reference than CHM13 for the genomic analysis of South Asian samples with less reference bias and better performance in mapping and variant calling, particularly for long read sequencing data. As the first high-quality T2T diploid reference genome for Indian, the largest worlds population, I002C contributes to the growing set of population-specific reference genomes and helps to overcome a significant gap in human genome diversity.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Centromeric transposable elements and epigenetic status drive karyotypic variation in the eastern hoolock gibbon 96%
- Genomic Insights into the Demographic History and Local Adaptation of Wild Boars Across Eurasia 96%
- Polymorphic short tandem repeats make widespread contributions to blood and serum traits 95%
Similar papers in this journal
- Genotyping sequence-resolved copy number variationusing pangenomes reveals paralog-specific global diversityand expression divergence of duplicated genes 96%
- Systematic assessment of regulatory effects of human disease variants in pluripotent cells 96%
- Prioritization of autoimmune disease-associated genetic variants that perturb regulatory element activity in T cells 95%
Similar papers in this journal
- A comprehensive catalog of 3D genome organization in diverse human genomes facilitates understanding of the impact of structural variation on chromatin structure 97%
- Prioritization of enhancer mutations by combining allele-specific chromatin accessibility with deep learning 96%
- Characterising tandem repeat complexities across long-read sequencing platforms with TREAT and otter 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.