Back

Scalable and interpretable secretion system annotation with Sismis

Larralde, M.; Albrecht, F.; Blom, J.; Henriksson, J.; Carroll, L. M.

2025-09-09 microbiology
10.1101/2025.09.09.675188 bioRxiv
Show abstract

Secretion systems play critical roles in bacterial growth, survival, and pathogenesis. Genes encoding secretion system components often co-occur together as gene clusters in bacterial (meta)genomes. However, existing tools for secretion system annotation are unable to utilize genomic context to make predictions and are unable to detect secretion systems of novel architecture. Here, we present Sismis (secretion system discovery tool; https://github.com/lmc297/Sismis), a scalable, interpretable, machine learning-based tool, which detects and classifies single-locus secretion systems in bacterial (meta)genomes with high accuracy (test set area-under-the-curve [AUC] values of 0.71 and 0.92 for precision-recall [PR] and receiver operating characteristic [ROC] curves, respectively). When applied to {approx}700k prokaryotic (meta)genomes, Sismis identifies 747, 439 total secretion systems comprising 15, 612 major secretion system families, >80% of which contain no previously known/annotated secretion systems. To facilitate further exploration of these data, we present the Sismis Atlas (https://sismis.microbe.dev/), an interactive secretion system database, which we use to identify a largely uncharacterized cluster of secretion systems with tight adherence (Tad) pili-like characteristics. Altogether, Sismis and its companion atlas enable accurate and interpretable secretion system annotation, exploration, and discovery at an unprecedented scale. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=77 SRC="FIGDIR/small/675188v1_ufig1.gif" ALT="Figure 1"> View larger version (23K): org.highwire.dtl.DTLVardef@89dec2org.highwire.dtl.DTLVardef@17fa652org.highwire.dtl.DTLVardef@18064b9org.highwire.dtl.DTLVardef@54deaa_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.