VILOCA: Sequencing quality-aware haplotype reconstruction and mutation calling for short- and long-read data
Fuhrmann, L.; Langer, B.; Topolsky, I.; Beerenwinkel, N.
Show abstract
RNA viruses exist in large heterogeneous populations within their host. The structure and diversity of virus populations affects disease progression and treatment outcomes. Next-generation sequencing allows detailed viral population analysis, but inferring diversity from error-prone reads is challenging. Here, we present VILOCA, a method for mutation calling and reconstruction of local haplotypes from short- and long-read viral sequencing data. Local haplotypes refer to genomic regions that have approximately the length of the input reads. VILOCA recovers local haplotypes by using a Dirichlet process mixture model to cluster reads around their unobserved haplotypes and leveraging quality scores of the sequencing reads. We assessed the performance of VILOCA in terms of mutation calling and haplotype reconstruction accuracy on simulated and experimental Illumina, PacBio, and Oxford Nanopore data. On simulated and experimental Illumina data, VILOCA performed better or similar to existing methods. On the simulated long-read data, VILOCA is able to recover on average 82% of the ground truth mutations with perfect precision compared to only 64% recall and 90% precision of the second-best method. In summary, VILOCA provides significantly improved accuracy in mutation and haplotype calling, especially for long-read sequencing data, and therefore facilitates the comprehensive characterization of heterogeneous within-host viral populations.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- An assembly-free method of phylogeny reconstruction using short-read sequences from pooled samples without barcodes 96%
- Demonstrating the utility of flexible sequence queries against indexed short reads with FlexTyper 95%
- Learning, Visualizing and Exploring 16S rRNA Structure Using an Attention-based Deep Neural Network 95%
Similar papers in this journal
- Vulcan: Improved long-read mapping and structural variant calling via dual-mode alignment 96%
- Pangenome databases provide superior host removal and mycobacteria classification from clinical metagenomic data 96%
- ntsm: an alignment-free, ultra low coverage, sequencing technology agnostic, intraspecies sample comparison tool for sample swap detection 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.