LncPNdeep: A long non-coding RNA classifier based on Large Language Model with peptide and nucleotide embedding
Dai, Z.; Deng, F.
Show abstract
Long non-coding RNA plays an important role in various gene transcription and peptide interactions. Classifying lncRNAs from coding RNA is a crucial step in bioinformatics analysis which seriously affects the post-analysis for transcriptome annotation. Although several machine learning-based methods were developed to classify lncRNAs, these methods were mainly focused on nucleotide features without considering the information from the peptide sequence. To integrate both nucleotide and peptide information in lncRNA classification, one efficient deep learning is desired. In this study, we developed one concatenated deep neural network named LncPNdeep to combine this information. LncPNdeep incorporates both peptide and nucleotide embedding from masked language modeling (MLM), being able to discover complex associations between sequence information and lncRNA classification. LncPNdeep achieves state-of-the-art performance in the human transcript database compared with other existing methods (Accuracy=97.1%). It also exhibits superior generalization ability in cross-species comparison, maintaining consistent accuracy and F1 scores compared to other methods. The combination of nucleotide and peptide information makes LncPNdeep able to facilitate the identification of novel lncRNA and gain high accuracy for classification. Our code is available at https://github.com/yatoka233/LncPNdeep
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeepLncLoc: a deep learning framework for long non-coding RNA subcellular localization prediction based on subsequence embedding 97%
- CRISPR-DIPOFF: An Interpretable Deep LearningApproach for CRISPR Cas-9 Off-Target Prediction 96%
- LSTM-PHV: Prediction of human-virus protein-protein interactions by LSTM with word2vec 96%
Similar papers in this journal
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 96%
- Representation learning applications in biological sequence analysis 95%
- Gra-CRC-miRTar: The pre-trained nucleotide-to-graph neural networks to identify potential miRNA targets in colorectal cancer 95%
Similar papers in this journal
- UTRGAN: Learning to Generate 5' UTR Sequences for Optimized Translation Efficiency and Gene Expression 95%
- LinAliFold and CentroidLinAliFold: Fast RNA consensus secondary structure prediction for aligned sequences using beam search methods 95%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.