Back

Shifu: an integrated framework for deep learning of RNA secondary structure

Galvez, G. C.; Vicens, Q.

2026-07-16 bioinformatics
10.64898/2026.07.15.738762 bioRxiv
Show abstract

Deep learning has advanced RNA secondary-structure prediction by bypassing explicit energy rules to capture long-range dependencies, yet progress is limited less by model scale than by how structures are measured: single scores hide where and why models fail, and benchmark scores can reflect memorization of one dataset rather than genuine generalization. We address this with Shifu, a framework of three coupled parts. Shifu-Corpus is a leakage-audited dataset of 254123 sequences from six databases, with family-aware splits certified free of exact and near-duplicate leaks. The Shifu Trifecta scores a model on three axes (correctness, breadth across diverse RNAs, and whether its confidence can be trusted) rather than one number. Shifu-LMR, a family of compact RNA language models, serves as controlled experiments: changing the training corpus shifts accuracy by 0.13, and a 65-million-parameter model, Shifu-LMR-Nano, leads on correctness while running on a laptop. We release the dataset, code, and model backbones.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.