Back

A refined characterization of large-scale genomic differences in the first complete human genome

Yang, X.; Wang, X.; Zou, Y.; Zhang, S.; Xia, M.; Vollger, M. R.; Chen, N.-C.; Taylor, D. J.; Harvey, W. T.; Logsdon, G. A.; Meng, D.; Shi, J.; McCoy, R. C.; Schatz, M. C.; Li, W.; Eichler, E. E.; Lu, Q.; Mao, Y.

2022-12-19 genomics
10.1101/2022.12.17.520860 bioRxiv
Show abstract

The first telomere-to-telomere (T2T) human genome assembly (T2T-CHM13) release was a milestone in human genomics. The T2T-CHM13 genome assembly extends our understanding of telomeres, centromeres, segmental duplication, and other complex regions. The current human genome reference (GRCh38) has been widely used in various human genomic studies. However, the large-scale genomic differences between these two important genome assemblies are not characterized in detail yet. Here, we identify 590 discrepant regions ([~]226 Mbp) in total. In addition to the previously reported non-syntenic regions, we identify 67 additional large-scale discrepant regions and precisely categorize them into four structural types with a newly developed website tool (SynPlotter). The discrepant regions ([~]20.4 Mbp) excluding telomeric and centromeric regions are highly structurally polymorphic in humans, where copy number variation are likely associated with various human disease and disease susceptibility, such as immune and neurodevelopmental disorders. The analyses of a newly identified discrepant region--the KLRC gene cluster--shows that the depletion of KLRC2 by a single deletion event is associated with natural killer cell differentiation in [~]20% of humans. Meanwhile, the rapid amino acid replacements within KLRC3 is consistent with the action of natural selection during primate evolution. Our study furthers our understanding of the large-scale structural variation differences between these two crucial human reference genomes and future interpretation of studies of human genetic variation.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.