Back

CluSeek: Bioinformatics Tool to Identify and Analyze Gene Clusters

Hrebicek, O.; Kadlcik, S.; Najmanova, L.; Janata, J.; Kamanova, J.; Hanzlikova, L.; Koberska, M.; Kovarovic, V.; Kamenik, Z.

2025-09-18 bioinformatics
10.1101/2025.09.16.676505 bioRxiv
Show abstract

Gene clusters are key structural and functional units that encode diverse phenotypes in genomes, from metabolism to pathogenesis. As genome sequencing expands, tools for systematic exploration of this growing data are increasingly needed. We present CluSeek, an open-source, Python-based platform for discovering, visualizing, and analyzing gene clusters across all GenBank data. Unlike existing tools, CluSeek does not rely on predefined cluster types or reference libraries but enables mining of any gene neighborhoods containing colocalized homologs of user-specified genes. It features an intuitive graphical interface suitable for non-bioinformaticians and is freely available at https://cluseek.com. We demonstrate the versatility of Cluseek in two distinct case studies: (i) mining of specialized metabolites, where CluSeek uncovered over 16 new classes containing the bioactivity-enhancing 4-alkyl-L-proline moiety, previously known in only three Golden Era antibiotic classes; and (ii) analysis of type III secretion systems present in Bordetella species, revealing previously unrecognized taxonomic distribution, and genetic variants, including gene multiplications and novel components with potential functional significance. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=135 SRC="FIGDIR/small/676505v1_ufig1.gif" ALT="Figure 1"> View larger version (30K): org.highwire.dtl.DTLVardef@1767e2eorg.highwire.dtl.DTLVardef@5603e3org.highwire.dtl.DTLVardef@1193475org.highwire.dtl.DTLVardef@1c32e83_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.