Scan Cluster: A versatile database-independent prediction tool for multi-genome identification of homologous gene clusters.
Mogro, E. G.; Pagnutti, A. L.; Zapata, G.; Lozano, M. J.
Show abstract
The exploration of colocalized gene sets, such as Biosynthetic Gene Clusters (BGCs) and symbiotic islands, is fundamental in modern genome mining. However, many existing prediction tools rely heavily on curated databases or predefined rules, inherently biasing detection toward known clusters. To address these limitations, we introduce Scan Cluster, a robust and flexible Python-based bioinformatic tool designed to identify user-defined, conserved gene clusters across diverse genomes without database-driven constraints. Scan Cluster leverages BLAST or HMMER to detect homologous proteins and evaluates their colocalization, accommodating complex evolutionary events such as orientation inversions, the insertion of alien genes, gene deletions, and the integration of insertion sequences. Beyond identification, the software performs progressive multiple cluster alignments and generates distance trees to assess cluster similarity, producing outputs ready for visualization in iTOL and Clinker. We validated Scan Cluster against standard tools like antiSMASH and DeepBGC, demonstrating high accuracy in delineating complete cluster boundaries. Its utility was further confirmed through the analysis of the nos and complex nod-nif-fix symbiotic gene clusters in rhizobia strains, successfully tracking genetic decay and grouping diverse architectures. Scan Cluster provides an accessible, low-resource framework to explore the evolutionary and functional dynamics of novel genetic clusters.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- MicrobeAnnotator: a user-friendly, comprehensive microbial genome annotation pipeline 96%
- PoMeLo: a systematic computational approach to predicting metabolic loss in pathogen genomes 95%
- ChiMera: An easy to use pipeline for Bacterial Genome Based Metabolic Network Reconstruction, Evaluation and Visualization 94%
Similar papers in this journal
- ExplorePipolin: reconstruction and annotation of bacterial mobile elements from draft genomes 97%
- MerCat2: a versatile k-mer counter and diversity estimator for database-independent property analysis obtained from omics data 95%
- BugBuster: A novel automatic and reproducible workflow for metagenomic data analysis 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.