Compression and k-mer based Approach For Anticancer Peptide Analysis
Ali, S.; Ali, T. E.; Chourasia, P.; Patterson, M.
Show abstract
Our research delves into the imperative realm of anti-cancer peptide sequence analysis, an essential domain for biological researchers. Presently, neural network-based methodologies, while exhibiting precision, encounter challenges with a substantial parameter count and extensive data requirements. The recently proposed method to compute the pairwise distance between the sequences using the compression-based approach [26] focuses on compressing entire sequences, potentially overlooking intricate neighboring information for individual characters (i.e., amino acids in the case of protein and nucleotide in the case of nucleotide) within a sequence. The importance of neighboring information lies in its ability to provide context and enhance understanding at a finer level within the sequences being analyzed. Our study advocates an innovative paradigm, where we integrate classical compression algorithms, such as Gzip, with a pioneering k-mersbased strategy in an incremental fashion. Diverging from conventional techniques, our method entails compressing individual k-mers and incrementally constructing the compression for subsequences, ensuring more careful consideration of neighboring information for each character. Our proposed method improves classification performance without necessitating custom features or pre-trained models. Our approach unifies compression, Normalized Compression Distance, and k-mers-based techniques to generate embeddings, which are then used for classification. This synergy facilitates a nuanced understanding of cancer sequences, surpassing state-of-the-art methods in predictive accuracy on the Anti-Cancer Peptides dataset. Moreover, our methodology provides a practical and efficient alternative to computationally demanding Deep Neural Networks (DNNs), proving effective even in low-resource environments.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PRIEST - Predicting viral mutations with immune escape capability of SARS-CoV-2 using temporal evolutionary information 96%
- Interpretable and Generalizable Attention-Based Model for Predicting Drug-Target Interaction Using 3D Structure of Protein Binding Sites: SARS-CoV-2 Case Study and in-Lab Validation 96%
- CRISPR-DIPOFF: An Interpretable Deep LearningApproach for CRISPR Cas-9 Off-Target Prediction 95%
Similar papers in this journal
- Pred-AHCP: Robust feature selection enabled Sequence Specific Prediction of Anti-Hepatitis C Peptides via Machine Learning 95%
- When does additional information improve accuracy of RNA secondary structure prediction? 94%
- Pathway-guided deep neural network toward interpretable and predictive modeling of drug sensitivity 94%
Similar papers in this journal
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 96%
- Predicting RNA Sequence-Structure Likelihood via Structure-Aware Deep Learning 95%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.