REINDEER2: practical abundance index at scale
Hernandez-Courbevoie, Y.; Salson, M.; Bessiere, C.; Yue, H.; Gautheret, D.; Marchet, C.; Limasset, A.
Show abstract
Recent advances in biological sequence indexing have enabled the efficient querying of sequence presence across massive genomic data repositories. While presence queries have become tractable at petabyte scale, retrieving quantitative information such as sequence abundances remains a significant algorithmic challenge. Existing abundance-aware indexes are mostly static, difficult to scale, and often trade off completeness, precision, or updatability. We describe a novel discrete abundance index designed for scalability, dynamic updates, and tunable precision. We combine an inverted index with probabilistic and exact structures to support fast, memory-efficient construction and precise high-throughput queries across thousands of RNA datasets. Our experiments demonstrate that our method REINDEER2 achieves one to two orders of magnitude speedup in construction compared to existing methods, while maintaining comparable or better memory use. Despite using approximate structures for scalability, REINDEER2 achieves sub-1% error on abundance recovery and correlates strongly with reference quantifiers like Kallisto. It also supports sequence-level queries in seconds over thousands of datasets. Code and experiments github.com/Yohan-HernandezCourbevoie/REINDEER2
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- kmtricks: Efficient and flexible construction of Bloom filters for large sequencing data collections 98%
- FroM Superstring to Indexing: a space-efficient index for unconstrained k-mer sets using the Masked Burrows-Wheeler Transform (MBWT) 98%
- K2R: Tinted de Bruijn Graphs implementation for efficient read extraction from sequencing datasets 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.