Back

DLRNA-BERTa: A transformer approach for RNA-drug binding affinity prediction

Lobascio, P.; Saeed, K.; Khan, A.; Tanoli, Z.

2025-09-06 molecular biology
10.1101/2025.09.05.674445 bioRxiv
Show abstract

RNA-based therapies are a rapidly expanding field, offering treatments for a wide range of diseases, including many rare conditions. To date, 24 RNA therapeutics have received FDA approval, with 131 more in clinical trials, underscoring RNAs growing role in modern medicine. In this context, Bidirectional Encoder Representations from Transformers (BERT) models provide a cost-effective and accurate virtual screening strategy for accelerating RNA-targeted drug discovery. These models take RNA FASTA sequences and compound SMILES strings as inputs and generate predicted binding affinities in nanomolar units. In this study, we introduce DLRNA-BERTa, a RoBERTa-based framework combining RNA-BERTa, pretrained on 9.76 million RNA sequences, with ChemBERTa-v2 for predicting small molecule-RNA interactions. The framework includes six class-specific models, aptamers, repeats, ribosomal RNAs, riboswitches, microRNAs (miRNAs), and viral RNAs, plus a general model for cases where the RNA class is unknown. Proposed DLRNA-BERTa consistently outperforms existing RNA-drug interaction prediction methods. Pearson correlation coefficients achieved are: 0.94 (aptamers), 0.95 (repeats), 0.93 (ribosomal RNAs), 0.94 (riboswitches), 0.95 (viral RNAs), 0.98 (miRNAs), and 0.92 (general model), demonstrating robust performance across RNA classes. Benchmarking against four independent datasets from the ROBIN repository further confirms generalizability. Application of DLRNA-BERTa to 3,492 approved drugs from the ChEMBL database identified 2,859 compounds with predicted affinities (pKd [≥] 6) across 294 RNA targets. As proof of concept, bleomycin is highlighted, supported by literature evidence of RNA-binding activity. A publicly accessible web application is available at https://huggingface.co/spaces/IlPakoZ/DLRNA-BERTa, in alignment with FAIR principles.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.