Stochastic Sampling of Structural Contexts Improves the Scalability and Accuracy of RNA 3D Module Identification
Sarrazin-Gendron, R.; Yao, H.-T.; Reinharz, V.; Oliver, C. G.; Ponty, Y.; Waldispuhl, J.
Show abstract
RNA structures possess multiple levels of structural organization. Secondary structures are made of canonical (i.e. Watson-Crick and Wobble) helices, connected by loops whose local conformations are critical determinants of global 3D architectures. Such local 3D structures consist of conserved sets of non-canonical base pairs, called RNA modules. Their prediction from sequence data is thus a milestone toward 3D structure modelling. Unfortunately, the computational efficiency and scope of the current 3D module identification methods are too limited yet to benefit from all the knowledge accumulated in modules databases. Here, we introduce BayesPairing 2, a new sequence search algorithm leveraging secondary structure tree decomposition which allows to reduce the computational complexity and improve predictions on new sequences. We benchmarked our methods on 75 modules and 6380 RNA sequences, and report accuracies that are comparable to the state of the art, with considerable running time improvements. When identifying 200 modules on a single sequence, BayesPairing 2 is over 100 times faster than its previous version, opening new doors for genome-wide applications.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep learning models for RNA secondary structure prediction (probably) do not generalise across families 98%
- Concurrent prediction of RNA secondary structures with pseudoknots and local 3D motifs in an Integer Programming framework 98%
- Graph neural representational learning of RNA secondary structures for predicting RNA-protein interactions 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- RNA-EFM : Energy based Flow Matching for Protein-conditioned RNA Sequence-Structure Co-design 95%
- LinAliFold and CentroidLinAliFold: Fast RNA consensus secondary structure prediction for aligned sequences using beam search methods 95%
- Prediction of RNA-protein interactions using a nucleotide language model 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.