RabbitTClust: enabling fast clustering analysis of millions bacteria genomes with MinHash sketches
Xu, X.; Yin, Z.; Yan, L.; Zhang, H.; Xu, B.; Wei, Y.; Niu, B.; Schmidt, B.; Liu, W.
Show abstract
We present RabbitTClust, a fast and memory-efficient genome clustering tool based on sketch-based distance estimation. Our approach enables efficient processing of large-scale datasets by combining dimensionality reduction techniques with streaming and parallelization on modern multi-core platforms. 113,674 complete bacterial genome sequences (RefSeq: 455 GB in FASTA format) can be clustered within less than 6 minutes and 1,009,738 GenBank assembled bacterial genomes (4.0 TB in FASTA format) within only 34 minutes on a 128-core workstation. Our results further identify 1,269 repetitive genomes (identical nucleotide content) in RefSeq bacterial genomes.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Sequences Dimensionality-Reduction by K-mer Substring Space Sampling Enables Effective Resemblance- and Containment-Analysis for Large-Scale omics-data 97%
- MetaBinner: a high-performance and stand-alone ensemble binning method to recover individual genomes from complex microbial communities 97%
- KegAlign: Optimizing pairwise alignments with diagonal partitioning 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.