CPPred-sORF: Coding Potential Prediction of sORF based on non-AUG
Tong, X.; Hong, X.; Xie, J.; Liu, S.
Show abstract
In recent years, researchers have discovered thousands of sORFs that can encode micropeptides, and more and more discoveries that non-AUG codons can be used as translation initiation sites for these micropeptides. On the basis of our previous tool CPPred, we develop CPPred-sORF by adding two features and using non-AUG as the starting codon, which makes a comprehensive evaluation of sORF. The database of CPPred-sORF are constructed by small coding RNA and lncRNA as positive and negative data, respectively. Compared to the small coding RNAs and small ncRNAs, lncRNAs and small coding RNAs are less distinguishable. This is because the longer the sequences, the easier to include open reading frames. We find that the sensitivity, specificity and MCC value of CPPred-sORF on the independent testing set can reach 88.22%, 88.84% and 0.768, respectively, which shows much better prediction performance than the other methods.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Rcirc: an R package for circRNA analyses and visualization 95%
- Characterization of Human Dosage-Sensitive Transcription Factor Genes 93%
- Epigenetic regulator miRNA pattern differences among SARS-CoV, SARS-CoV-2 and SARS-CoV-2 world-wide isolates delineated the mystery behind the epic pathogenicity and distinct clinical characteristics of pandemic COVID-19 93%
Similar papers in this journal
- Network based multifactorial modelling of miRNA-target interactions 93%
- SC-JNMF: Single-cell clustering integrating multiple quantification methods based on joint non-negative matrix factorization 92%
- Construction of competing endogenous RNA interaction networks as prognostic markers in metastatic melanoma 92%
Similar papers in this journal
- DeepLncLoc: a deep learning framework for long non-coding RNA subcellular localization prediction based on subsequence embedding 96%
- An in silico approach to identification, categorization and prediction of nucleic acid binding proteins 96%
- Normalization of RNA-Seq Data using Adaptive Trimmed Mean with Multi-reference 95%
Similar papers in this journal
- bpRNA-align: Improved RNA Secondary Structure Global Alignment for Comparing and Clustering RNA Structures 91%
- Unraveling Unbreakable Hairpins: Characterizing RNA secondary structures that are persistent after dinucleotide shuffling 91%
- Evaluating DCA-based method performances for RNA contact prediction by a well-curated dataset 90%
Similar papers in this journal
- SubFeat: Feature Subspacing Ensemble Classifier for Function Prediction of DNA, RNA and Protein Sequences 92%
- Exploring vulnerable building blocks in protein-protein interaction networks of breast tumor and adjacent normal tissues 91%
- A Nile Grass Rat Transcriptomic Landscape Across 22 Organs By Ultra-deep Sequencing and Comparative RNA-seq pipeline (CRSP) 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.