Back

nf-core/viralmetagenome: A Novel Pipeline for Untargeted Viral Genome Reconstruction

Klaps, J.; Lemey, P.; nf-core community, ; Kafetzopoulou, L. E.

2025-07-02 bioinformatics
10.1101/2025.06.27.661954 bioRxiv
Show abstract

MotivationEukaryotic viruses present significant challenges for genome reconstruction and variant analysis due to their extensive diversity and potential genome segmentation. While de novo assembly followed by reference database matching and scaffolding is a commonly used approach, the manual execution of this workflow is extremely time-consuming, particularly due to the extensive reference curation required. Here, we address the critical need for an automated, scalable pipeline that can efficiently handle viral metagenomic analysis without manual intervention. ResultsWe present nf-core/viralmetagenome, a comprehensive viral metagenomic pipeline for untargeted genome reconstruction and variant analysis of eukaryotic DNA and RNA viruses. Viral-metagenome is implemented as a Nextflow workflow that processes short-read metagenomic samples to automatically detect and assemble viral genomes, while also performing variant analysis. The pipeline features automated reference selection, consensus quality control metrics, comprehensive documentation, and seamless integration with containerization technologies, including Docker and Singularity. We demonstrate the utility and accuracy of our approach through validation on both simulated and real datasets, showing robust performance across diverse viral families in metage-nomic samples. Availabilitynf-core/viralmetagenome is freely available at https://github.com/nf-core/viralmetagenome with comprehensive documentation at https://nf-co.re/viralmetagenome Contactjoon.klaps@kuleuven.be Supplementary informationSupplementary data are available at https://github.com/Joon-Klaps/nf-core-viralmetagenome-manuscript online.

Published in Bioinformatics (predicted rank #5) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.