Back

An RNA foundation model enables discovery of disease mechanisms and candidate therapeutics

Celaj, A.; Gao, A. J.; Lau, T. T. Y.; Holgersen, E. M.; Lo, A.; Lodaya, V.; Cole, C. B.; Denroche, R. E.; Spickett, C.; Wagih, O.; Pinheiro, P. O.; Vora, P.; Mohammadi-Shemirani, P.; Chan, S.; Nussbaum, Z.; Zhang, X.; Zhu, H.; Ramamurthy, E.; Kanuparthi, B.; Iacocca, M.; Ly, D.; Kron, K.; Verby, M.; Cheung-Ong, K.; Shalev, Z.; Vaz, B.; Bhargava, S.; Yusuf, F.; Samuel, S.; Alibai, S.; Baghestani, Z.; He, X.; Krastel, K.; Oladapo, O.; Mohan, A.; Shanavas, A.; Bugno, M.; Bogojeski, J.; Schmitges, F.; Kim, C.; Grant, S.; Jayaraman, R.; Masud, T.; Deshwar, A.; Gandhi, S.; Frey, B. J.

2023-09-26 bioinformatics
10.1101/2023.09.20.558508 bioRxiv
Show abstract

Accurately modeling and predicting RNA biology has been a long-standing challenge, bearing significant clinical ramifications for variant interpretation and the formulation of tailored therapeutics. We describe a foundation model for RNA biology, "BigRNA", which was trained on thousands of genome-matched datasets to predict tissue-specific RNA expression, splicing, microRNA sites, and RNA binding protein specificity from DNA sequence. Unlike approaches that are restricted to missense variants, BigRNA can identify pathogenic non-coding variant effects across diverse mechanisms, including polyadenylation, exon skipping and intron retention. BigRNA accurately predicted the effects of steric blocking oligonucleotides (SBOs) on increasing the expression of 4 out of 4 genes, and on splicing for 18 out of 18 exons across 14 genes, including those involved in Wilson disease and spinal muscular atrophy. We anticipate that BigRNA and foundation models like it will have widespread applications in the field of personalized RNA therapeutics.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.