Back

Sparse and Skew Hashing of K-Mers

Pibiri, G. E.

2022-01-18 bioinformatics
10.1101/2022.01.15.476199 bioRxiv
Show abstract

MotivationA dictionary of k-mers is a data structure that stores a set of n distinct k-mers and supports membership queries. This data structure is at the hearth of many important tasks in computational biology. High-throughput sequencing of DNA can produce very large k-mer sets, in the size of billions of strings - in such cases, the memory consumption and query efficiency of the data structure is a concrete challenge. ResultsTo tackle this problem, we describe a compressed and associative dictionary for k-mers, that is: a data structure where strings are represented in compact form and each of them is associated to a unique integer identifier in the range [0, n). We show that some statistical properties of k-mer minimizers can be exploited by minimal perfect hashing to substantially improve the space/time trade-off of the dictionary compared to the best-known solutions. AvailabilityThe C++ implementation of the dictionary is available at https://github.com/jermp/sshash. Contactgiulio.ermanno.pibiri@isti.cnr.it

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 0.7%
27.8%
2
Algorithms for Molecular Biology
17 papers in training set
Top 0.1%
19.4%
3
Genome Research
468 papers in training set
Top 0.7%
7.0%
50% of probability mass above
4
PLOS Computational Biology
1863 papers in training set
Top 8%
4.5%
5
Journal of Computational Biology
48 papers in training set
Top 0.3%
3.4%
6
PLOS ONE
5266 papers in training set
Top 39%
2.9%
7
Peer Community Journal
281 papers in training set
Top 2%
2.8%
8
Nature Biotechnology
172 papers in training set
Top 2%
2.8%
9
iScience
1154 papers in training set
Top 10%
2.6%
10
Nature Communications
5641 papers in training set
Top 39%
2.6%
11
PeerJ
308 papers in training set
Top 3%
2.5%
12
Genome Biology
637 papers in training set
Top 5%
2.2%
13
Bioinformatics Advances
203 papers in training set
Top 3%
2.0%
14
Cell Systems
201 papers in training set
Top 2%
1.8%
15
BMC Bioinformatics
457 papers in training set
Top 4%
1.4%
16
Nature Methods
385 papers in training set
Top 5%
1.2%
17
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 0.7%
1.2%
18
Scientific Reports
3612 papers in training set
Top 71%
1.0%
19
IEEE Transactions on Computational Biology and Bioinformatics
20 papers in training set
Top 0.7%
0.6%
20
NAR Genomics and Bioinformatics
242 papers in training set
Top 5%
0.6%
21
GigaScience
212 papers in training set
Top 5%
0.6%
22
Nature Computational Science
55 papers in training set
Top 2%
0.5%
23
Molecular Biology and Evolution
542 papers in training set
Top 6%
0.5%