Shifu: an integrated framework for deep learning of RNA secondary structure
Galvez, G. C.; Vicens, Q.
Show abstract
Deep learning has advanced RNA secondary-structure prediction by bypassing explicit energy rules to capture long-range dependencies, yet progress is limited less by model scale than by how structures are measured: single scores hide where and why models fail, and benchmark scores can reflect memorization of one dataset rather than genuine generalization. We address this with Shifu, a framework of three coupled parts. Shifu-Corpus is a leakage-audited dataset of 254123 sequences from six databases, with family-aware splits certified free of exact and near-duplicate leaks. The Shifu Trifecta scores a model on three axes (correctness, breadth across diverse RNAs, and whether its confidence can be trusted) rather than one number. Shifu-LMR, a family of compact RNA language models, serves as controlled experiments: changing the training corpus shifts accuracy by 0.13, and a 65-million-parameter model, Shifu-LMR-Nano, leads on correctness while running on a laptop. We release the dataset, code, and model backbones.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Single-sequence protein structure prediction using language models from deep learning 95%
- Human 5′ UTR design and variant effect prediction from a massively parallel translation assay 95%
- Multi-omics integration and regulatory inference for unpaired single-cell data with a graph-linked unified embedding framework 94%
Similar papers in this journal
Similar papers in this journal
- Measuring intramolecular connectivity in long RNA molecules using two-dimensional DNA patch-probe arrays 95%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 95%
- A Generalizable Scaffold-Based Approach for Structure Determination of RNAs by Cryo-EM 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.