Back

Discovering motifs and genomic patterns with SMT: a high-performance data structure for counting kmers

Caldonazzo Garbelini, J. M.; Sanches, D. S.; Pozo, A. T. R.

2023-04-02 bioinformatics
10.1101/2023.04.01.535163 bioRxiv
Show abstract

MotivationThe search for conserved motifs in DNA sequences is an important problem in bioinformatics. The growing availability of large-scale genomic data poses significant challenges for computational biology, particularly in terms of efficiency in analysis, kmer identification, and noise presence. The detection of conserved motifs and patterns in DNA sequences is crucial for understanding gene functions and regulations. Therefore, it is essential to develop a data structure that can handle these large volumes of information and provide accurate and fast results. ResultsWe present SMT, an innovative tool designed to efficiently store and count kmers, optimizing memory usage and computation time. The application of SMT has also proven effective in discovering motifs in noisy datasets, allowing the identification of conserved regions in sequences. Furthermore, SMT enables exact searches in constant time and recovers the most abundant k-mers, as well as performs approximate searches in linear time to find fragments with up to d mutations. This approach facilitates large-scale data analysis and provides important insights into the conserved properties of biological sequences. The application of SMT in motif discovery demonstrates its potential to drive research in bioinformatics and genomics. Supplementary data and results are available to provide additional information and support the conclusions presented in this work. Availability and implementationThe source code of the presented method is publicly available at https://github.com/jadermcg/SMT.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.