vOMIX-MEGA: An ultra-fast end-to-end pipeline for terabyte-scale viral metagenomics analysis.
SHEKARRIZ, E.; VIJENDRAN, E.; Ho, J. W. K.
Show abstract
Viral identification for terabyte-scale metagenomic data is limited by scalability and computational resources. We present vOMIX-MEGA, an end-to-end viral metagenomic framework that overcomes performance bottlenecks by significantly improving parallelization and memory usage in critical steps. Benchmarked on empirical datasets, it completes processing in up to 50 minutes with 24 GB of RAM, bypassing four other state-of-the-art pipelines that require 7 hours (383 GB) to 14 days (32 GB). vOMIX-MEGA is on average 21% and 13% more accurate when benchmarked on mock and experimental data and is available via https://github.com/holab-hku/vOMIX-MEGA.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- nf-core/viralmetagenome: A Novel Pipeline for Untargeted Viral Genome Reconstruction 97%
- V-pipe: a computational pipeline for assessing viral genetic diversity from high-throughput sequencing data 94%
- TaxTriage: An Open-Source Metagenomic Sequencing Data Analysis Pipeline Enabling Putative Pathogen Detection 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.