A fully-open structure-guided RNA foundation model for robust structural and functional inference
Zhu, H.; Li, R.; Zhang, F.; Tang, F.; Ye, T.; Li, X.; Gu, Y.; Xiong, P.; Zhou, S. K.
Show abstract
RNA language models have achieved strong performances across diverse down-stream tasks by leveraging large-scale sequence data. However, RNA function is fundamentally shaped by its hierarchical structure, making the integration of structural information into pre-training essential. Existing methods often depend on noisy structural annotations or introduce task-specific biases, limiting model generalizability. Here, we introduce structRFM, a structure-guided RNA foundation model that is pre-trained on millions of RNA sequences and secondary structures data by integrating base pairing interactions into masked language modeling through a novel pair matching operation. The structure-guided mask and nucleotide-level mask are further balanced by a dynamic masking ratio. structRFM learns joint knowledge of sequential and structural data, producing versatile representations, including classification-level, sequence-level, and pair-wise matrix features, that support a broad spectrum of downstream adaptations. structRFM ranks among the top models in zero-shot homology classification across fifteen biological language models, and sets new benchmarks for secondary structure prediction. structRFM further derives Zfold, which enables robust and reliable tertiary structure prediction, with consistent wimprovements in estimating 3D structures and their accordingly extracted 2D structures, achieving a pronounced 19% performance gain compared with AlphaFold3 on RNA Puzzles dataset. In functional tasks such as internal ribosome entry site identification, structRFM achieves a whopping 49% performance gain in F1 score. These results demonstrate the effectiveness of structure-guided pre-training and highlight a promising direction for developing multi-modal RNA language models in computational biology. To support the broader scientific community, we have made the 21-million sequence-structure dataset and the pre-trained structRFM model fully open-source, facilitating the development of multimodal foundation models in biology.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A 5' UTR Language Model for Decoding Untranslated Regions of mRNA and Function Predictions 97%
- Predicting RNA 3D structure and conformers using a pre-trained secondary structure model and structure-aware attention 95%
- Construction of a 3D whole organism spatial atlas by joint modeling of multiple slices 94%
Similar papers in this journal
- Developing a general AI model for integrating diverse genomic modalities and comprehensive genomic knowledge 95%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 95%
- scIGANs: single-cell RNA-seq imputation using generative adversarial networks 94%
Similar papers in this journal
- JEDI: Circular RNA Prediction based on Junction Encoders and Deep Interaction among Splice Sites 95%
- DeepLocRNA: An Interpretable Deep Learning Model for Predicting RNA Subcellular Localization with domain-specific transfer-learning 95%
- ProtNote: a multimodal method for protein-function annotation 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.