Back

SparseRNAFolD: Sparse RNA pseudoknot-free Folding including Dangles

Gray, M.; Jabbari, H.; Will, S.

2023-06-07 bioinformatics
10.1101/2023.06.05.543808 bioRxiv
Show abstract

MotivationComputational RNA secondary structure prediction by free energy minimization is indispensable for analyzing structural RNAs and their interactions. These methods find the structure with the minimum free energy (MFE) among exponentially many possible structures and have a restrictive time and space complexity (O(n3) time and O(n2) space for pseudoknot-free structures) for longer RNA sequences. Furthermore, accurate free energy calculations including dangles contributions can be difficult and costly to implement, particularly when optimizing for time and space requirements. ResultsHere we introduce a fast and efficient sparsified MFE pseudoknot-free structure prediction algorithm, SparseRNAFolD, that utilizes an accurate energy model that accounts for dangles contributions. While sparsification technique was previously employed to improve time and space complexity of a pseudoknot-free structure prediction method with a realistic energy model, SparseMFEFold, it was not extended to include dangles contributions due to complexity of computation. This may be at the cost of prediction accuracy. In this work, we compare three different sparsified implementations for dangles contributions and provide pros and cons of each method. As well, we compare our algorithm to LinearFold, a linear time and space algorithm, where we find in practice, SparseRNAFolD has lower memory consumption across all lengths of sequence and a faster time for lengths up to 1000 bases. ConclusionOur SparseRNAFolD algorithm is an MFE-based algorithm that guarantees optimality of result and employs the most general energy model including dangles contributions. We provide basis for applying dangles to sparsified recursion in a pseudoknot-free model which has the ability to be extended to pseudoknots. AvailabilitySparseRNAFolDs algorithm and detailed results are available at https://github.com/mateog4712/SparseRNAFolD.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.