RNAGEN: A generative adversarial network-based model to generate synthetic RNA sequences to target proteins
Ozden, F.; Barazandeh, S.; Akboga, D.; Seker, U. O. S.; Cicek, A. E.
Show abstract
RNA - protein binding plays an important role in regulating protein activity by affecting localization and stability. While proteins are usually targeted via small molecules or other proteins, easy-to-design and synthesize small RNAs are a rather unexplored and promising venue. The problem is the lack of methods to generate RNA molecules that have the potential to bind to certain proteins. Here, we propose a method based on generative adversarial networks (GAN) that learn to generate short RNA sequences with natural RNA-like properties such as secondary structure and free energy. Using an optimization technique, we fine-tune these sequences to have them bind to a target protein. We use RNA-protein binding prediction models from the literature to guide the model. We show that even if there is no available guide model trained specifically for the target protein, we can use models trained for similar proteins, such as proteins from the same family, to successfully generate a binding RNA molecule to the target protein. Using this approach, we generated piRNAs that are tailored to bind to SOX2 protein using models trained for its relative (SOX10, SOX14, and SOX8) and experimentally validated in vitro that the top-2 molecules we generated specifically bind to SOX2.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- UTRGAN: Learning to Generate 5' UTR Sequences for Optimized Translation Efficiency and Gene Expression 97%
- LinAliFold and CentroidLinAliFold: Fast RNA consensus secondary structure prediction for aligned sequences using beam search methods 95%
- Prediction of RNA-protein interactions using a nucleotide language model 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.