Back

Accurate haplotype-resolved de novo assembly of human genomes with RFhap

Gonzalez, D.; Cabas, G.; Miquel, J. F.; Moraga, C.; Salas, F.; Di Genova, A.

2026-01-30 bioinformatics
10.64898/2026.01.28.702238 bioRxiv
Show abstract

Haplotype-resolved de novo assemblies enable genome-wide separation of maternal and paternal variation, improving the interpretation of complex variants relevant to human disease. Trio-aware assemblers such as Hifiasm-Trio leverage parental short reads by deriving parent-specific k-mers to guide phasing within the long-read assembly graph; however, fixed k-mer length and heuristics cannot be optimal in the presence of sequencing errors and graph complexity in repetitive regions, contributing to phasing errors and reduced haplotype-resolved contiguity. Here we present RFhap, a trio-based long-read phasing method that integrates multi-k-mer parent-specific markers with an alignment-free k-mer lookup engine and a random forest classifier to assign long reads to maternal, paternal, or unknown haplotypes prior to de novo assembly. We benchmarked RFhap on four human trio datasets from the Human Pangenome project spanning two ONT chemistries (R10.4 and R9.4.1). Using Merqury to evaluate downstream assemblies, RFhap nearly doubled corrected haplotype NG50 (mean 24.3 Mb vs 13.1 Mb) and halved switch error rates (mean 0.111% vs 0.236%) relative to Hifiasm-Trio, while remaining competitive in consensus QV and parental k-mer completeness. Consistent with improved phasing, RFhap reduced long-switch errors by [~]3-fold across datasets, particularly within interspersed repeats. Together, these results demonstrate that RFhap improves phasing accuracy from standard trio data, representing a step toward accurate and automated diploid human de novo assembly from long reads.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.