Back

Prediction of lncRNA-protein interacting pairs using LLM embeddings based on evolutionary information

Choudhury, S.; Bajiya, N.; Raghava, G. P. S.

2025-11-11 bioinformatics
10.1101/2025.11.10.687101 bioRxiv
Show abstract

Interactions of long non-coding RNAs (lncRNAs) with proteins is responsible for numerous cellular processes, including transcriptional regulation, chromatin remodeling, cell differentiation, and intracellular signaling. In the past, numerous computational methods have been developed for predicting lncRNA-protein interacting (LPI) pairs. This study describes a highly accurate and reliable method for predicting LPI pairs built on largest possible non-redundant dataset having 262,244 interacting and equal number of non-interacting pairs. Initially, similarity-based approach BLAST has been tried which have poor discriminative power, due to low sequence similarity. Subsequently, we developed CNN based models and machine learning based models using traditional features and embedding. Our CatBoost model developed using embedding generated by DNABERT-2 and ESM-2-t30 achieved AUC of 0.989 with MCC 0.915 on an independent dataset. Our method performs better than existing methods on an independent dataset. We developed standalone software and web server lncrnaPI for predicting LPI pairs, scanning lncRNA interacting proteins in proteome and protein interacting lncRNA in genomes (https://webs.iiitd.edu.in/raghava/lncrnapi/). HIGHLIGHTSO_LIDiscrimination of LncRNA-protein interacting and non-interacting pairs. C_LIO_LINon-redundant dataset of 262,244 interacting and 262,244 non-interacting pairs. C_LIO_LIEmbedding of LncRNA using DNABERT-2 and protein using ESM-2-t30. C_LIO_LIRapid scanning of lncRNA interacting proteins at genome scale. C_LIO_LIA web server and software for predicting LncRNA-protein interacting pairs. C_LI

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.