Back

Rapid and Consistent Genome Clustering for Navigating Bacterial Diversity with Millions of MAGs and Isolates

von Wachsmann, J. H.; Lorenz, L. J.; Gurbich, T.; Russell, M.; Rodriguez Bouza, V.; Horsfield, S.; Lees, J. A.; Finn, R. D.

2025-12-30 bioinformatics
10.64898/2025.12.30.695181 bioRxiv
Show abstract

Bacterial genome databases now exceed 7 million assemblies; however, the massive redundancy and limited scalability of existing tools create bottlenecks for large-scale analyses. Current clustering methods struggle beyond a few thousand genomes, making database-wide organisation computationally infeasible. Here, we present gemsparcl, a tool that clusters bacterial genomes at species-level resolution over 400x faster than existing methods. We developed sketchlib.rust, implementing one-permutation MinHash with densification in binned sketches to accelerate all-versus-all comparisons. Combined with distance correction of incomplete metagenome-assembled genomes (MAGs) for accurate distance estimation and network-based quality filtering to edges, gemsparcl clusters genomes into biologically coherent species-level groups. We clustered 2.2 million bacterial genomes (1.86 million isolates and 360,000 MAGs) into 15,837 species-level genomic cohesive units (GCUs) in 12 hours using 32 cores and less than 64GB of memory. The method achieves 99.8% species purity on high-quality isolates while maintaining accuracy across mixed-quality datasets. This advance enables routine database maintenance for resources like MGnify, as well as reference-free microbiome analysis across millions of genomes, and database-wide metagenomics studies that were previously impossible due to computational constraints. Gemsparcl transforms bacterial genome organisation from a months-long challenge requiring high-performance computing into an overnight analysis on standard hardware.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.