Back

vOMIX-MEGA: An ultra-fast end-to-end pipeline for terabyte-scale viral metagenomics analysis.

SHEKARRIZ, E.; VIJENDRAN, E.; Ho, J. W. K.

2026-07-31 bioinformatics
10.64898/2026.07.28.741255 bioRxiv
Show abstract

Viral identification for terabyte-scale metagenomic data is limited by scalability and computational resources. We present vOMIX-MEGA, an end-to-end viral metagenomic framework that overcomes performance bottlenecks by significantly improving parallelization and memory usage in critical steps. Benchmarked on empirical datasets, it completes processing in up to 50 minutes with 24 GB of RAM, bypassing four other state-of-the-art pipelines that require 7 hours (383 GB) to 14 days (32 GB). vOMIX-MEGA is on average 21% and 13% more accurate when benchmarked on mock and experimental data and is available via https://github.com/holab-hku/vOMIX-MEGA.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.