HydraRNA: a hybrid architecture based full-length RNA language model
Li, G.; Jiang, F.; Zhu, J.; Cui, H.; Wang, Z.; Chen, W.
Show abstract
RNA, an essential component of the central dogma of molecular biology, plays versatile roles in all cellular processes. RNA large language models (LLMs) are emerging as powerful methods in RNA research to decipher its intricate network of function and regulation. However, previous RNA LLMs were based on the transformer model and pre-trained on short segment of non-coding RNAs, which limits their general usability. Here we present the first full-length RNA foundation model, HydraRNA, which is based on a hybrid architecture of bidirectional state space model and multi-head attention mechanism, and is pre-trained on a large amount of both protein-coding and non-coding RNAs. Despite being pre-trained with the fewest parameters and the least GPU resources, HydraRNA learns better RNA representations and outperforms the existing foundation models on a variety of downstream tasks, including RNA classification, prediction of RNA secondary structure, RBP binding sites, mRNA stability and translation efficiency. Furthermore, HydraRNA can accurately predict the effect of mutations and estimate the relative contributions of different mRNA regions to the RNA stability and translation. We anticipate that HydraRNA will enable dissecting the diverse properties of RNA, accelerating the research of RNA regulation and facilitating the optimal design of RNA therapeutics.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ERNIE-RNA: An RNA Language Model with Structure-enhanced Representations 99%
- Benchmarking Pre-trained Genomic Language Models for RNA Sequence-Related Predictive Applications 98%
- Differential Analysis of RNA Structure Probing Experiments at Nucleotide Resolution: Uncovering Regulatory Functions of RNA Structure 97%
Similar papers in this journal
- DeepCLIP: Predicting the effect of mutations on protein-RNA binding with Deep Learning 96%
- ModiDeC: a multi-RNA modification classifier for direct nanopore sequencing 95%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 94%
Similar papers in this journal
- US-align: Universal Structure Alignments of Proteins, Nucleic Acids, and Macromolecular Complexes 96%
- A systematic benchmark of Nanopore long read RNA sequencing for transcript level analysis in human cell lines 95%
- Absolute quantitative and base-resolution sequencing reveals comprehensive landscape of pseudouridine across the human transcriptome 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.