PanSpace: Fast and Scalable Indexing for Massive Bacterial Databases
Avila Cartes, J. E.; Ciccolella, S.; Denti, L.; Raghuram Dandinasivara, R.; Della Vedova, G.; Bonizzoni, P.; Schonhuth, A.
Show abstract
MotivationSpecies identification is a critical task in agriculture, food processing, and health-care. The rapid growth of genomic databases -- driven in part by the increasing investigation of bacterial genomes in clinical microbiology -- has outpaced the capabilities of conventional tools such as BLAST for basic search and query tasks. A key bottleneck in microbiome studies lies in building indexes that allow rapid species identification and classification from assemblies while scaling efficiently to massive resources such as the AllTheBacteria database, thus enabling large-scale analyses to be performed even on a common laptop. ResultsWe introduce PanSpace, the first convolutional neural network-based approach that leverages dense vector (embedding) indexing --- scalable to billions of embeddings --- for indexing and querying massive bacterial genome databases. PanSpace is specifically designed to classify bacterial draft assemblies. Compared to the most recent and competitive tool for this task, PanSpace requires only ~2 GB of disk space to index the AllTheBacteria database, an 8x reduction relative to existing methods. Moreover, it delivers ultra-fast query performance, processing more than 1,000 assemblies in less than two and a half minutes, while preserving the utmost accuracy of state-of-the-art approaches. AvailabilityPanSpace is available at https://github.com/pg-space/panspace.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- KPop: Accurate and scalable comparative analysis of microbial genomes by sequence embeddings 97%
- Simplitigs as an efficient and scalable representation of de Bruijn graphs 96%
- Sequences Dimensionality-Reduction by K-mer Substring Space Sampling Enables Effective Resemblance- and Containment-Analysis for Large-Scale omics-data 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.