Back

Sassy: Searching Short DNA Strings in the 2020s

Beeloo, R.; Koerkamp, R. G.

2025-07-26 bioinformatics
10.1101/2025.07.22.666207 bioRxiv
Show abstract

MotivationApproximate string matching (ASM) is the problem of finding all occurrences of a pattern in a text while allowing up to k errors. Many modern methods use seed-chain-extend, which is fast in practice, but does not guarantee finding all matches with [≤] k errors. However, applications such as CRISPR off-target detection require exhaustive results. MethodsWe introduce Sassy, a library and tool for ASM of short patterns in long texts. Sassy splits the text into 4 parts that are searched in parallel, and uses bitvectors in the text direction rather than the pattern direction. This has compexity O(k{lceil}n/W {rciel}) when searching a random text of length n, where W = 256 is the SIMD width, and provides significant speedups for small k. Separately, we allow matches of the pattern to extend beyond the text for an overhang cost of e.g. = 0.5 per character, to find matches near contig or read ends. ResultsSassy is 4x to 15x faster than Edlib for patterns [≤] 1000bp, and can search text with a throughput near 2 Gbp/s. Likewise, Sassy is over 100x faster than parasail. We apply Sassy to CRISPR off-target detection by searching 61 guide sequences in a human genome. Sassy is 100x faster than SWOffinder and only slightly slower (for k [≤] 3) than CHOPOFF, for which building its index takes 20 minutes. Sassy also scales well to larger k, unlike CHOPOFF whose index took over 10 hours to build for k = 5. AvailibilitySassy is available as library and binary at https://github.com/RagnarGrootKoerkamp/sassy, and archived at swh:1:dir:e884758dce5777a441bc2799dc8824e563c5f97b.

Published in Bioinformatics (predicted rank #1) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.