Back

Viral haplotype reconstruction from long reads with virCHap

Gao, Y.; Yu, T.; Liu, B.; Li, G.

2026-02-10 bioinformatics
10.64898/2026.02.09.704753 bioRxiv
Show abstract

Resolving genomes at the haplotype level for viral populations is crucial for understanding the prevalence of viral diseases and for the development of effective therapeutic treatments. However, viral haplotype reconstruction still presents challenges, such as an unknown number of strains, high inter-strain similarity, repetitive regions, and difficulties in abundance estimation. Here, we developed virCHap, a new reference-based haplotype phasing algorithm for viruses, which applies graph partitioning followed by iteratively quantifiable cluster merging on long-read sequencing data. Benchmarking on simulated and real datasets demonstrates that virCHap outperforms current tools in terms of recall, accurate abundance estimates and read clustering accuracy. On the simulated large-genome VZV experiment, virCHap had a 96.5% recall, 14% higher than the second-best method, and had the most accurate abundance estimates. On a real 5-strain PVY dataset, virCHap had a precision exceeding 92.9%, a recall of over 97%, and a read clustering accuracy of 82%, outperforming the second-best method by 32%. On a real 6-strain SARS-CoV-2 dataset, virCHap achieved >96.9% accuracy, and the most accurate abundance estimates within the spike gene.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.