Back

Language models enable zero-shot prediction of RNA secondary structures including pseudoknots

Gong, T.; Bu, D.

2024-01-29 bioinformatics
10.1101/2024.01.27.577533 bioRxiv
Show abstract

Current deep learning-based models for predicting RNA secondary structures face challenges in achieving high generalization ability. At the same time, a vast repository of unlabeled non-coding RNA (ncRNA) sequences remains untapped for structure prediction tasks. To address this challenge, we trained RNA-km, a foundation language model that enables zero-shot prediction of RNA secondary structures including pseudoknots. For the end, we incorporated specific modifications into the language model training process, including k-mer masking strategy and relative positional encoding. RNA-km are trained on 23 million ncRNA sequences in a self-supervised manner, gaining the advantages of high generalization ability. For a target RNA sequence, we make a zero-shot secondary structure prediction with the attention maps provided by RNA-km and a specified minimum-cost flow algorithm. Our results on popular benchmark datasets demonstrate that RNA-km exhibits high generalization abilities, excelling in zero-shot predictions for RNA secondary structures. In addition, the attention maps provided by the model capture intricate structural relationships, as evidenced by accurate pseudoknot predictions and precise identification of longdistance base pairs. We anticipate that RNA-km enhances the predictive capacity and robustness of existing models, thereby improving their ability to accurately predict structures for novel RNA sequences.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.