Back

FoldARE, an RNA secondary structure analysis and prediction tool via generative pseudo-SHAPE modeling

Marino, S. M.; Husak, V.; Tebaldi, T.

2026-03-05 bioinformatics
10.64898/2026.03.04.709501 bioRxiv
Show abstract

RNA secondary structure prediction remains limited by RNA conformational heterogeneity and scarcity of experimental data: many RNAs populate ensembles of near-isoenergetic folds, and structure-probing data such as SHAPE are often unavailable. Here we introduce FoldARE (Folding and Analysis of RNA Ensembles), a two-step framework that derives pseudo-SHAPE constraints from in silico structural ensembles and uses them to guide a downstream SHAPE-aware predictor. In the first ("ensembler") step, an ensemble is generated and parsed position-by-position to estimate single-strandedness frequencies, which are converted into a pseudo-SHAPE reactivity profile via a weight-and-threshold scheme. In the second ("predictor") step, this profile is supplied as a constraint to a SHAPE-compatible folding algorithm to produce an improved secondary structure. We systematically evaluated all combinations of four ensemble-capable predictors: ViennaRNA, RNAstructure, LinearFold, and EternaFold. We optimized parameters on a manually curated, structurally diverse 25-RNA training set and validated robustness using multiple scoring schemes, including identity-based measures and an exact base-pair-matching metric. The optimal configuration uses EternaFold as the ensembler and RNAstructure as the predictor, yielding consistent gains over all standalone methods and over other ensembler/predictor pairings. On external benchmark RNA structure collections (RNAstrand, ArchiveII, and bpRNA; total n = 1964 after filtering) and on the experimentally derived eFold dataset spanning human mRNAs, pre-miRNAs, and lncRNAs (n = 1024), FoldARE achieved the highest accuracy across datasets with highly significant improvements in paired comparisons. Beyond prediction, FoldARE provides modules for ensemble-level comparative analysis, including pairwise and multi-tool consensus assessment, per-nucleotide variability metrics, and interactive visualizations. It also supports the evaluation of m6A modification effects on folding ensembles using modification-aware engines. Together, our results show that mining ensemble statistics to generate pseudo-probing constraints is an effective, accessible strategy to improve RNA secondary structure prediction and to support ensemble-focused structural analysis. FoldARE is freely available on GitHub (https://github.com/TebaldiLab/FoldARE).

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.