Comprehensive benchmarking of tools to identify phages in metagenomic shotgun sequencing data
Ho, S. F. S.; Millard, A. D.; van Schaik, W.
Show abstract
BackgroundThe prediction of bacteriophage sequences in metagenomic datasets has become a topic of considerable interest, leading to the development of many novel bioinformatic tools. A comparative analysis of ten state-of-the-art phage identification tools was performed to inform their usage in microbiome research. MethodsArtificial contigs generated from complete RefSeq genomes representing phages, plasmids, and chromosomes, and a previously sequenced mock community containing four phage species, were used to evaluate the precision, recall and F1-scores of the tools. We also generated a dataset of randomly shuffled sequences to quantify false positive calls. In addition, a set of previously simulated viromes was used to assess diversity bias in each tools output. ResultsVirSorter2 achieved the highest F1 score (0.92) in the RefSeq artificial contigs dataset, with several other tools also performing well. Kraken2 had the highest F1 score (0.86) in the mock community benchmark by a large margin (0.3 higher than DeepVirFinder in second place), mainly due to its high precision (0.96). Generally, k-mer based tools performed better than reference similarity tools and gene-based methods. Several tools, most notably PPR Meta, called a high number of false positives in the randomly shuffled sequences. When analysing the diversity of the genomes that each tool predicted from a virome set, most tools produced a viral genome set that had similar alpha and beta diversity patterns to the original population, with Seeker being a notable exception. ConclusionsThis study provides key metrics used to assess performance of phage detection tools, offers a framework for further comparison of additional viral discovery tools, and discusses optimal strategies for using these tools. We highlight that the choice of tool for identification of phages in metagenomic datasets, as well as their parameters, can bias the results and provide pointers for different use case scenarios. We have also made our benchmarking dataset available for download in order to facilitate future comparisons of phage identification tools.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Nanopore and Illumina Sequencing Reveal Different Viral Populations from Human Gut Samples 96%
- From defaults to databases: parameter and database choice dramatically impact the performance of metagenomic taxonomic classification tools 95%
- FANGORN: A quality-checked and publicly available database of full-length 16S-ITS-23S rRNA operon sequences 95%
Similar papers in this journal
Similar papers in this journal
- Addressing the dynamic nature of reference data: a new nt database for robust metagenomic classification 95%
- Long-term incubation of lake water enables genomic sampling of consortia involving Planctomycetes and Candidate Phyla Radiation bacteria 94%
- Clade-specific long-read sequencing increases the accuracy and specificity of the gyrB phylogenetic marker gene 94%
Similar papers in this journal
- IDseq - An Open Source Cloud-based Pipeline and Analysis Service for Metagenomic Pathogen Detection and Monitoring 97%
- dadasnake, a Snakemake implementation of DADA2 to process amplicon sequencing data for microbial ecology 96%
- PathoGFAIR: a collection of FAIR and adaptable (meta)genomics workflows for (foodborne) pathogens detection and tracking 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.