Comparative evaluation of computational methods for reconstruction of human viral genomes
Sousa, M. J. P.; Toppinen, M.; Pyöriä, L.; Hedman, K.; Sajantila, A.; Perdomo, M. F.; Pratas, D.
Show abstract
The increasing availability of viral sequences has led to the emergence of many optimized viral genome reconstruction tools. Given that the number of new tools is steadily increasing, it is complex to identify functional and optimized tools that offer an equilibrium between accuracy and computational resources as well as the features that each tool provides. In this paper, we surveyed open-source computational tools (including pipelines) used for human viral genome reconstruction, identifying specific characteristics, features, similarities, and dissimilarities between these tools. For quantitative comparison, we create an open-source reconstruction benchmark based on viral data. The benchmark was executed using both synthetic and real datasets. With the former, we evaluated the effects to the reconstruction process of using different human viruses with simulated mutation rates, contamination and mitochondrial DNA inclusion, and various coverage depths. Each reconstruction program was also evaluated using real datasets, demonstrating their performance in real-life scenarios. The evaluation measures include the identity, a Normalized Compression Semi-Distance, and the Normalized Relative Compression between the genomes before and after reconstruction, as well as metrics regarding the length of the genomes reconstructed, computational time and resources spent by each tool. The benchmark is fully reproducible and freely available at https://github.com/viromelab/HVRS.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Real-time resolution of short-read assembly graph using ONT long reads 96%
- AvP: a software package for automatic phylogenetic detection of candidate horizontal gene transfers. 94%
- Variant calling tool evaluation for variable size indel calling from next generation whole genome and targeted sequencing data 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.