Redesigning the Eterna100 for the Vienna 2 folding engine
Koodli, R. V.; Rudolfs, B.; Wayment-Steele, H. K.; Eterna Structure Designers, ; Das, R.
Show abstract
The rational design of RNA is becoming important for rapidly developing technologies in medicine and biochemistry, spurring development of numerous RNA secondary structure design algorithms and benchmarks to evaluate their performance. However, the problem of RNA design is dependent upon the reverse problem of RNA structure prediction through "folding engines" that predict structure from sequence. We hypothesized that differences in RNA folding engines could impact design algorithms, and recruited an online community of RNA design experts to modify the widely-used RNA secondary structure design benchmark, Eterna100, to address unsolvability of some cases when changing the folding engine used (Vienna 1.8 updated to Vienna 2.6). We tested this new Eterna100-V2 benchmark with five RNA design algorithms, and found that while overall rankings remained similar, the performance of RNA design algorithms that depended on folding engines in their training did indeed depend on which underlying parameter set was used in training. This work demonstrates that the design "difficulty" of RNA structures is intrinsically linked to thermodynamic models, and suggests that future RNA design algorithms that are agnostic to thermodynamic models will result in optimal performance and development. Eterna100-V1 and Eterna100-V2 benchmarks and example solutions are freely available at https://github.com/eternagame/eterna100-benchmarking. Author SummaryDesigning RNA molecules that fold to a desired target structure is an algorithmic problem gathering increasing attention due to the emergence of RNA-based therapies and the need for rational design of RNAs. The Eterna100 dataset, a collection of target structures with increasing design difficulty, designed and selected by players of the crowdsourced RNA game Eterna, has been widely used to benchmark RNA design algorithms. However, these puzzles were originally developed using the now-deprecated version 1 of the ViennaRNA folding engine. In this manuscript, we introduce an updated benchmark, called the Eterna100-V2. We found that nineteen puzzles using Vienna 1 were unsolvable in Vienna 2, but that Eterna participants were able to re-design the puzzles with minimal modifications to make them solvable in Vienna 2. We confirmed that the rankings of 5 RNA design algorithms remained consistent between Eterna100-V1 and -V2. However, discrepancies in performance from algorithms that relied on thermodynamic models in their training suggest that algorithms will benefit from being agnostic to thermodynamic models as these models continue to improve.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep learning models for RNA secondary structure prediction (probably) do not generalise across families 98%
- Concurrent prediction of RNA secondary structures with pseudoknots and local 3D motifs in an Integer Programming framework 96%
- RNAtive to recognize native-like structure in a set of RNA 3D models 96%
Similar papers in this journal
- bpRNA-align: Improved RNA Secondary Structure Global Alignment for Comparing and Clustering RNA Structures 97%
- Evaluating DCA-based method performances for RNA contact prediction by a well-curated dataset 95%
- Unraveling Unbreakable Hairpins: Characterizing RNA secondary structures that are persistent after dinucleotide shuffling 95%
Similar papers in this journal
Similar papers in this journal
- LinAliFold and CentroidLinAliFold: Fast RNA consensus secondary structure prediction for aligned sequences using beam search methods 96%
- Estimating Protein Complex Model Accuracy Using Graph Transformers and Pairwise Similarity Graphs 94%
- RNA-EFM : Energy based Flow Matching for Protein-conditioned RNA Sequence-Structure Co-design 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.