Back

Efficient and accurate near telomere-to-telomere haplotype reconstruction of diploid genomes

Liu, Y.; Yichen, L.; Xu, J.; Tan, Z.; Zhang, W.; Wang, L.; Xu, L.; Zeng, X.; Schoenhuth, A.; Luo, X.

2026-05-24 bioinformatics
10.64898/2026.05.20.726711 bioRxiv
Show abstract

Telomere-to-telomere (T2T) and haplotype-resolved assembly are crucial for understanding eukaryotic genomes. For diploid species, this resolution is critical to uncover allelic variations, inheritance patterns, and functional genomic traits. Current scaffolding methods typically employ either sequence-based or graph-based strategies. Sequence-based approaches rely on proximity signals to yield high contiguity, but underutilize assembly graph information, resulting in more structural errors and chromosomal misassignments. Graph-based methods leverage graph topology for higher accuracy but frequently struggle to achieve chromosome-scale contiguity. However, neither strategy alone can overcome its inherent limitations to simultaneously achieve high contiguity and accuracy. To address these challenges, we introduce HapFold, the first hybrid scaffolding framework that synergistically leverages the complementary strengths of both graph-based and sequence-based approaches. By integrating the topological accuracy of assembly graphs with the proximity-guided contiguity of sequence models, HapFold achieves highly accurate, chromosome-scale or near-T2T haplotype reconstructions for diploid genomes. Compared to existing methods, HapFold achieves superior assembly quality while accelerating computation by an order of magnitude. Furthermore, in the haplotype reconstruction of diploid genomes using standard Oxford Nanopore Technologies simplex reads, HapFold enables the reconstruction of a greater number of near-T2T assemblies. Our approach provides a robust and scalable solution for the high-fidelity reconstruction of haplotype-resolved diploid genomes.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Genome Research
468 papers in training set
Top 0.1%
18.1%
2
Genome Biology
637 papers in training set
Top 0.8%
11.6%
3
Nucleic Acids Research
1281 papers in training set
Top 2%
9.6%
4
Bioinformatics
1204 papers in training set
Top 3%
9.5%
5
Nature Communications
5641 papers in training set
Top 22%
7.7%
50% of probability mass above
6
Nature Biotechnology
172 papers in training set
Top 0.4%
7.1%
7
Nature Methods
385 papers in training set
Top 2%
4.7%
8
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
3.1%
9
Briefings in Bioinformatics
354 papers in training set
Top 3%
2.7%
10
Cell Systems
201 papers in training set
Top 2%
1.9%
11
PLOS Computational Biology
1863 papers in training set
Top 14%
1.9%
12
Nature Computational Science
55 papers in training set
Top 0.8%
1.4%
13
GENETICS
483 papers in training set
Top 4%
1.1%
14
GigaScience
212 papers in training set
Top 4%
1.1%
15
Nature
645 papers in training set
Top 9%
1.1%
16
BMC Genomics
406 papers in training set
Top 6%
1.1%
17
BMC Bioinformatics
457 papers in training set
Top 5%
1.0%
18
Journal of Computational Biology
48 papers in training set
Top 1%
0.8%
19
G3: Genes, Genomes, Genetics
252 papers in training set
Top 4%
0.8%
20
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 42%
0.8%
21
Advanced Science
286 papers in training set
Top 10%
0.8%
22
Bioinformatics Advances
203 papers in training set
Top 5%
0.8%
23
Scientific Reports
3612 papers in training set
Top 80%
0.6%
24
Nature Genetics
286 papers in training set
Top 6%
0.6%