Back

EnhancerP-2L: A Gene regulatory site identification tool for DNA enhancer region using CREs motifs

Butt, A. H.; Alkhalaf, S.; Iqbal, S.; Khan, Y. D.

2020-01-20 bioinformatics
10.1101/2020.01.20.912451 bioRxiv
Show abstract

Enhancers are DNA fragments that do not encode RNA molecules and proteins, but they act critically in the production of RNAs and proteins by controlling gene expression. Prediction of enhancers and their strength plays significant role in regulating gene expression. Prediction of enhancer regions, in sequences of DNA, is considered a difficult task due to the fact that they are not close to the target gene, have less common motifs and are mostly tissue/cell specific. In recent past, several bioinformatics tools were developed to discriminate enhancers from other regulatory elements and to identify their strengths as well. However the need for improvement in the quality of its prediction method requires enhancements in its application value practically. In this study, we proposed a new method that builds on nucleotide composition and statistical moment based features to distinguish between enhancers and non-enhancers and additionally determine their strength. Our proposed method achieved accuracy better than current state-of-the-art methods using 5-fold and 10-fold cross-validation. The outcomes from our proposed method suggest that the use of statistical moments based features could bear more efficient and effective results. For the accessibility of the scientific community, we have developed a user-friendly web server for EnhancerP-2L which will increase the impact of bioinformatics on medicinal chemistry and drive medical science into an unprecedented resolution. Web server is freely accessible at http://www.biopred.org/enpred.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.