ViroSeek: a viral detection pipeline for second-generation sequencing
Berger, A.; Lefebvre, M. J. M.; Dainat, J.; Jiolle, D.; Conclois, I.; Talignani, L.; Mastriani, E.; Cornelie, S.; Berthet, N.; Paupy, C.
Show abstract
Arbovirus emergences represent a rising public health issue and are exacerbated by climate change and globalization. Virome analysis has become a key approach for monitoring and managing infectious diseases, yet existing tools often remain technically complex and inaccessible to non-specialists. In this context, we present ViroSeek, a lightweight, reproducible and accessible bioinformatics pipeline specifically designed for the taxonomic analysis of second-generation sequencing data. ViroSeek performs a series of automated steps: quality control (FastQC), trimming (TrimGalore or fastp), host and bacterial sequence removal (BBduk), assembly (SPAdes), taxonomic assignment (DIAMOND and TaxonKit), read remapping for quantification (minimap2), and PCR duplicate removal (Samtools). The whole process is designed to produce a clear, usable viral taxonomy table that is suitable for diversity studies. ViroSeek was empirically validated on enriched control samples containing a known panel of viruses. All the expected viruses were correctly detected. Bacterial and host contaminant sequences were effectively removed. The pipeline is freely available and fully documented, supporting its adoption and adaptation by the research community. It provides an optimized solution for virome analysis.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- mbctools: A User-Friendly Metabarcoding and Cross-Platform Pipeline for Analyzing Multiple Amplicon Sequencing Data across a Large Diversity of Organisms 93%
- Structural variation turnovers and defective genomes: key drivers for the in vitro evolution of the large double-stranded DNA koi herpesvirus (KHV) 93%
- EnVhogDB: an extended view of the viral protein families on Earth through a vast collection of HMM profiles 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.