Back

metaSMASH: Scalable Biosynthetic Gene Cluster Detection for Large Metagenomic Assemblies

Bagci, C.; Blin, K.; Ziemert, N.

2026-07-25 bioinformatics
10.64898/2026.07.22.739405 bioRxiv
Show abstract

antiSMASH is widely used for biosynthetic gene cluster (BGC) detection and annotation, but its standard workflow is poorly suited to large metagenomic assemblies, where massive contig counts create severe runtime bottlenecks and complicate downstream result exploration. We present metaSMASH, a re-engineered fork of antiSMASH for metagenome-scale BGC analysis. metaSMASH preserves the original antiSMASH detection and annotation logic while introducing streaming, memory-bounded execution, record-level parallelisation, optional output filtering, and an interactive dashboard for large result sets. Across 25 benchmark metagenome datasets, metaSMASH reproduced identical BGC detection results while dramatically reducing computational cost. Relative to the default antiSMASH configuration, metaSMASH was a geometric-mean 38x faster. It also outperformed an ad hoc chunked antiSMASH workflow: in the default configuration it achieved a geometric-mean 2.9 x speed-up and 1.7 x lower peak memory, and with extended-analysis modules enabled it was 2.7 x faster and used 3.1 x less memory while completing all datasets, whereas the ad hoc workflow ran out of memory on the two largest assemblies. By substantially reducing the computational burden of large-scale metagenome analysis without sacrificing result equivalence, metaSMASH makes routine mining of assembled metagenomes more practical and provides a scalable foundation for natural product discovery from complex microbial communities. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=78 SRC="FIGDIR/small/739405v1_ufig1.gif" ALT="Figure 1"> View larger version (27K): org.highwire.dtl.DTLVardef@d224ccorg.highwire.dtl.DTLVardef@6ddbdcorg.highwire.dtl.DTLVardef@7d5d40org.highwire.dtl.DTLVardef@7537e1_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.