Deciphering RNA regulation with a foundation language model
Zhou, H.; Hu, Y.; Zheng, Y.; Li, J.; Peng, J.; Zhang, G.; Wang, Z.
Show abstract
RNA metabolism is tightly regulated by cis-elements and trans-acting factors. Most information guiding such regulation is encoded in RNA sequences. Considering the similarities in semantic and syntactic features between RNAs and human language, we developed LAMAR, a transformer-based foundation language model for RNA regulation, to decipher general rules underlying RNA processing. The model was pretrained on approximately 15 million sequences from both genome and transcriptome of 225 mammals and 1569 viruses, and further fine-tuned with labeled datasets for various tasks. The resulting fine-tuned models outperformed the state-of-the-art methods in predicting mRNA translation efficiency and mRNA half-life, while achieving comparable accuracy to specifically designed methods in predicting splice sites of pre-mRNAs and internal ribosome entry sites. Our results indicated that a single foundation language model is applicable in the comprehensive analysis of different aspects of RNA regulation, providing new insight into the design and optimization of RNA drugs.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Differential Analysis of RNA Structure Probing Experiments at Nucleotide Resolution: Uncovering Regulatory Functions of RNA Structure 97%
- Benchmarking Pre-trained Genomic Language Models for RNA Sequence-Related Predictive Applications 97%
- ERNIE-RNA: An RNA Language Model with Structure-enhanced Representations 97%
Similar papers in this journal
- DEMINERS enables clinical metagenomics and comparative transcriptomic analysis by increasing throughput and accuracy of nanopore direct RNA sequencing 96%
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 96%
- Splam: a deep-learning-based splice site predictor that improves spliced alignments 95%
Similar papers in this journal
- DeepCLIP: Predicting the effect of mutations on protein-RNA binding with Deep Learning 96%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 96%
- ModiDeC: a multi-RNA modification classifier for direct nanopore sequencing 95%
Similar papers in this journal
- Absolute quantitative and base-resolution sequencing reveals comprehensive landscape of pseudouridine across the human transcriptome 96%
- A systematic benchmark of Nanopore long read RNA sequencing for transcript level analysis in human cell lines 96%
- Systematic assessment of long-read RNA-seq methods for transcript identification and quantification 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.