Back

Approaching an Error-Free Diploid Human Genome

Chu, Y.; Huang, Z.; Shao, C.; Guo, S.; Yu, X.; Wang, J.; Tian, Y.; Chen, J.; Li, R.; He, Y.; Yu, J.; Huang, J.; Gao, Z.; Kang, Y.

2025-08-01 genomics
10.1101/2025.08.01.667781 bioRxiv
Show abstract

SUMMARYAchieving an error-free diploid human genome remains challenging. We report T2T-YAO v2.0, a telomere-to-telomere complete assembly of a Han Chinese individual, polished to near-perfect base-level and structural accuracy. To systematically identify assembly errors, we developed Sufficient Alignment Support (SAS), an automatic method that flags structural and base-level errors in windows lacking sufficient read support. Building on this, we established a "structural-error-first" polishing strategy, correcting misassemblies using ultra-long ONT reads, followed by base-level refinement with PWC (Platform-integrated Window Consensus). Using these approaches, we resolved all detectable structure and non-homopolymer-related errors outside ribosomal DNA (rDNA) regions. The resulting assembly contains no unsupported 21-mers across sequencing platforms, meeting k-mer-based criteria for an error-free genome. T2T-YAO v2.0 delivers the most perfect East Asian reference to date, with limited issues confined to rDNA arrays and homopolymer tracks, enabling precise genome annotation, benchmarking, and variant discovery--foundation for human genomics and precision medicine.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.