Viral haplotype reconstruction from long reads with virCHap
Gao, Y.; Yu, T.; Liu, B.; Li, G.
Show abstract
Resolving genomes at the haplotype level for viral populations is crucial for understanding the prevalence of viral diseases and for the development of effective therapeutic treatments. However, viral haplotype reconstruction still presents challenges, such as an unknown number of strains, high inter-strain similarity, repetitive regions, and difficulties in abundance estimation. Here, we developed virCHap, a new reference-based haplotype phasing algorithm for viruses, which applies graph partitioning followed by iteratively quantifiable cluster merging on long-read sequencing data. Benchmarking on simulated and real datasets demonstrates that virCHap outperforms current tools in terms of recall, accurate abundance estimates and read clustering accuracy. On the simulated large-genome VZV experiment, virCHap had a 96.5% recall, 14% higher than the second-best method, and had the most accurate abundance estimates. On a real 5-strain PVY dataset, virCHap had a precision exceeding 92.9%, a recall of over 97%, and a read clustering accuracy of 82%, outperforming the second-best method by 32%. On a real 6-strain SARS-CoV-2 dataset, virCHap achieved >96.9% accuracy, and the most accurate abundance estimates within the spike gene.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Capturing variation in metagenomic assembly graphs with MetaCortex 96%
- V-pipe: a computational pipeline for assessing viral genetic diversity from high-throughput sequencing data 96%
- Demixer: A probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample 96%
Similar papers in this journal
Similar papers in this journal
- A Computational Toolset for Rapid Identification of SARS-CoV-2, other Viruses, and Microorganisms from Sequencing Data 95%
- Accurate identification of structural variations from cancer samples 94%
- VIGA: an one-stop tool for eukaryotic Virus Identification and Genome Assembly from next-generation-sequencing data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.