RIFT-VAE: grammar-conditioned pretraining and latent-space optimization for RNA inverse folding
Watanabe, K.; Akiyama, M.; Sakakibara, Y.
Show abstract
BackgroundRNA inverse folding designs nucleotide sequences expected to adopt a prescribed secondary structure. Search-based solvers can optimize folding-model objectives effectively, but difficult targets can require extensive sampling, and structural optimization alone does not explicitly preserve the sequence distributions or conserved motifs of natural RNA families. MethodsWe developed RIFT-VAE, a Transformer-based conditional variational autoencoder that receives a context-free grammar parse-tree representation of a target secondary structure and generates nucleotide labels on the corresponding tree. The framework combines progressively richer grammar rules, self-refinement learning from generated structure-sequence pairs, and cross-entropy-method optimization in the learned latent space. We evaluated RNAfold minimum-free-energy agreement on an RNAcentral-derived test set and the EteRNA100 benchmark, compared the method with four search-based solvers under matched total time budgets, and examined GC-content control and covariance-model family annotation. ResultsThe complete pipeline achieved RNAfold-Correct/RNAfold-MCC values of 0.833/0.994 on the RNAcentral-derived test set and 0.760/0.977 on EteRNA100. Latent-space optimization accounted for the largest increase in exact structural recovery. Under a 3,600-s total budget on EteRNA100, sequences generated by RIFT-VAE improved the exact-match rate of every tested downstream search method when used as warm starts; the largest change was observed for RNAInverse (Correct, 0.297 to 0.803; MCC, 0.505 to 0.985). The pretrained model also produced sequences with measurable correct-family covariance-model hits and supported explicit GC-content conditioning. ConclusionsRIFT-VAE is best interpreted as a hybrid generative-search framework: pretraining supplies a structure- and family-informed proposal distribution, whereas latent optimization concentrates evaluations in high-scoring regions. The reported structural scores are specific to RNAfold minimum-free-energy validation and do not establish biochemical function. Orthogonal folding predictors, stricter homology-controlled splits, diversity-aware evaluation, architecture-matched dot-bracket ablations, and experimental assays remain priorities for validation.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- RNAtive to recognize native-like structure in a set of RNA 3D models 95%
- DeepLocRNA: An Interpretable Deep Learning Model for Predicting RNA Subcellular Localization with domain-specific transfer-learning 94%
- DUETT quantitatively identifies known and novel events in nascent RNA structural dynamics from chemical probing data 94%
Similar papers in this journal
- Cat_Wiz: A stereochemistry-guided toolkit for locating, diagnosing and annotating Mg2+ ions in RNA structures. 95%
- IPANEMAP: Integrative Probing Analysis of Nucleic Acids Empowered by Multiple Accessibility Profiles 94%
- Accurate inference of the full base-pairing structure of RNA by deep mutational scanning and covariation-induced deviation of activity 94%
Similar papers in this journal
- EternaBrain: Automated RNA design through move sets from an Internet-scale RNA videogame 96%
- Global Importance Analysis: An Interpretability Method to Quantify Importance of Genomic Features in Deep Neural Networks 95%
- RNA structure prediction using positive and negative evolutionary information 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.