RNA secondary structure packages ranked and improved by high-throughput experiments
Wayment-Steele, H. K.; Kladwang, W.; Eterna Participants, ; Das, R.
Show abstract
The computer-aided study and design of RNA molecules is increasingly prevalent across a range of disciplines, yet little is known about the accuracy of commonly used structure modeling packages in tasks sensitive to ensemble properties of RNA. Here, we demonstrate that the EternaBench dataset, a set of over 20,000 synthetic RNA constructs designed in iterative cycles on the RNA design platform Eterna, provides incisive discriminative power in evaluating current packages in ensemble-oriented structure prediction tasks. We find that CONTRAfold and RNAsoft, packages with parameters derived through statistical learning, achieve consistently higher accuracy than more widely used packages in their standard settings, which derive parameters primarily from thermodynamic experiments. Motivated by these results, we develop a multitask-learning-based model, EternaFold, which demonstrates improved performance that generalizes to diverse external datasets, including complete mRNAs and viral genomes probed in human cells and synthetic designs modeling mRNA vaccines.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- US-align: Universal Structure Alignments of Proteins, Nucleic Acids, and Macromolecular Complexes 96%
- Direct prediction of intrinsically disordered protein conformational properties from sequence 94%
- Absolute quantitative and base-resolution sequencing reveals comprehensive landscape of pseudouridine across the human transcriptome 93%
Similar papers in this journal
- Protein Language Models Trained on Biophysical Dynamics Inform Mutation Effects 95%
- Parametrically guided design of beta barrels and transmembrane nanopores using deep learning 94%
- LinearTurboFold: Linear-Time Global Prediction of Conserved Structures for RNA Homologs with Applications to SARS-CoV-2 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.