Back

PLA-complexity of k-mer multisets

Abrar, M. H.; Medvedev, P.

2024-02-11 bioinformatics
10.1101/2024.02.08.579510 bioRxiv
Show abstract

MotivationUnderstanding structural properties of k-mer multisets is crucial to designing space-efficient indices to query them. A potentially novel source of structure can be found in the rank function of a k-mer multiset. In particular, the rank function of a k-mer multiset can be approximated by a piece-wise linear function with very few segments. Such an approximation was shown to speed up suffix array queries and sequence alignment. However, a more comprehensive study of the structure of rank functions of k-mer multisets and their potential applications is lacking. ResultsWe study a measure of a k-mer multiset complexity, which we call the PLA-complexity. The PLA-complexity is the number of segments necessary to approximate the rank function of a k-mer multiset with a piece-wise linear function so that the maximum error is bounded by a predefined threshold. We describe, implement, and evaluate the PLA-index, which is able to construct, compact, and query a piece-wise linear approximation of the k-mer rank function. We examine the PLA-complexity of more than 500 genome spectra and several other genomic multisets. Finally, we show how the PLA-index can be applied to several downstream applications to improve on existing methods: speeding up suffix array queries, decreasing the index memory of a short-read aligner, and decreasing the space of a direct access table of k-mer ranks. AvailabilityThe software and reproducibility information is freely available at https://github.com/medvedevgroup/pla-index

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 0.9%
22.7%
2
Algorithms for Molecular Biology
17 papers in training set
Top 0.1%
22.3%
3
Genome Research
468 papers in training set
Top 0.3%
12.1%
50% of probability mass above
4
PLOS Computational Biology
1863 papers in training set
Top 7%
5.2%
5
Journal of Computational Biology
48 papers in training set
Top 0.2%
4.4%
6
Cell Systems
201 papers in training set
Top 2%
2.8%
7
BMC Bioinformatics
457 papers in training set
Top 3%
2.4%
8
Genome Biology
637 papers in training set
Top 4%
2.4%
9
Nature Biotechnology
172 papers in training set
Top 2%
2.2%
10
PLOS ONE
5266 papers in training set
Top 44%
2.2%
11
iScience
1154 papers in training set
Top 12%
2.2%
12
PeerJ
308 papers in training set
Top 5%
1.8%
13
Bioinformatics Advances
203 papers in training set
Top 3%
1.8%
14
Nature Communications
5641 papers in training set
Top 44%
1.8%
15
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 0.6%
1.5%
16
Nature Methods
385 papers in training set
Top 5%
1.4%
17
Peer Community Journal
281 papers in training set
Top 5%
0.9%
18
Scientific Reports
3612 papers in training set
Top 73%
0.9%
19
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
0.9%
20
Nature Computational Science
55 papers in training set
Top 2%
0.6%
21
IEEE Transactions on Computational Biology and Bioinformatics
20 papers in training set
Top 0.7%
0.6%