RNAGym: Large-scale Benchmarks for RNA Fitness and Structure Prediction
Arora, R.; Angelo, M.; Choe, C. A.; Shearer, C.; Kollasch, A.; Qu, F.; Weitzman, R.; Gazizov, A.; Gurev, S.; Xie, E.; Marks, D. S.; Notin, P.
Show abstract
Understanding RNA structure and predicting the functional consequences of mutations are fundamental challenges in computational biology with broad implications for therapeutic development and synthetic biology. Current evaluation of machine learning-based RNA models suffers from disparate experimental datasets and inconsistent performance assessments across different RNA families. To address these challenges, we introduce RNAGym, a large-scale benchmarking framework specifically designed for three core tasks-RNA fitness, secondary structure, and tertiary structure prediction. The framework integrates extensive datasets, including 70 standardized deep mutational scanning assays covering over a million mutations across diverse RNA types; 901k chemical-mapping reactivity profiles for secondary structure; and 215 diverse tertiary structures curated from the PDB. RNAGym is designed to facilitate a systematic comparison of RNA models, offering an essential resource to enhance the understanding and development of these models.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep Learning for RNA Secondary Structure Determination: Gauging Generalizability and Broadening the Scope of Traditional Methods 97%
- Unspecific binding but specific disruption of the group I intron by the StpA chaperone 95%
- bpRNA-align: Improved RNA Secondary Structure Global Alignment for Comparing and Clustering RNA Structures 95%
Similar papers in this journal
- Deep learning models for RNA secondary structure prediction (probably) do not generalise across families 98%
- RNAtive to recognize native-like structure in a set of RNA 3D models 96%
- DUETT quantitatively identifies known and novel events in nascent RNA structural dynamics from chemical probing data 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.