Deep generalizable prediction of RNA secondary structure via base pair motif energy
Zhu, H.; Tang, F.; Quan, Q.; Chen, K.; Xiong, P.; Zhou, S. K.
Show abstract
Deep learning methods have demonstrated great performance for RNA secondary structure prediction. However, generalizability is a common unsolved issue on unseen out-of-distribution RNA families, which hinders further improvement of the accuracy and robustness of deep learning methods. Here we construct a base pair motif library that enumerates the complete space of the locally adjacent three-neighbor base pair and records the thermodynamic energy of corresponding base pair motifs through de novo modeling of tertiary structures, and we further develop a deep learning approach for RNA secondary structure prediction, named BPfold, which learns relationship between RNA sequence and the energy map of base pair motif. Experiments on sequence-wise and family-wise datasets have demonstrated the great superiority of BPfold compared to other state-of-the-art approaches in accuracy and generalizability. We hope this work contributes to integrating physical priors and deep learning methods for the further discovery of RNA structures and functionalities.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- sincFold: end-to-end learning of short- and long-range interactions in RNA secondary structure 96%
- A Reproducibility Analysis-based Statistical Framework for Residue-Residue Evolutionary Coupling Detection 96%
- KGETCDA: an efficient representation learning framework based on knowledge graph encoder from transformer for predicting circRNA-disease associations 95%
Similar papers in this journal
- ARTEMIS - a method for topology-independent superposition of RNA 3D structures and structure-based sequence alignment 95%
- UFold: Fast and Accurate RNA Secondary Structure Prediction with Deep Learning 95%
- A comprehensive survey of long-range tertiary interactions and motifs in non-coding RNA structures 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.