A Generalized Similarity Metric for Peptide Binding Affinity Prediction
Rodriguez, J. L.; Rath, S. S.; Francis-Landau, J. S.; Demirci, Y.; Ustundag, B. B.; Sarikaya, M.
Show abstract
The ability to capture the relationship between similarity and functionality would enable the predictive design of peptide sequences for a wide range of implementations from developing new drugs to molecular scaffolds in tissue engineering and biomolecular building blocks in nanobiotechnology. Similarity matrices are widely used for detecting sequence homology but depend on the assumption that amino acid mutational frequencies reflected by each matrix are relevant to the system in which they are applied. Increasingly, neural networks and other statistical learning models solve problems related to functional prediction but avoid using known features to circumvent unconscious bias. We demonstrated an iterative alignment method that enhances predictive power of similarity matrices based on a similarity metric, the Total Similarity Score. A generalized method is provided for application to amino acid sequences from inorganic and organic systems by benchmarking it on the debut quartz-binder set and 3 peptide-protein sets from the Immune Epitope Database. Pearson and Spearman Rank Correlations show that by treating the gapless Total Similarity Score as a predictor of relative binding affinity, prediction of test data has a 0.5-0.7 Pearson and Spearman Rank correlation. with respect to size of the dataset. Since the benchmarks used herein are from a solid-binding peptide and a protein-peptide system, our proposed method could prove to be a highly effective general approach for establishing the predictive sequence-function relationships of among the peptides with different sequences and lengths in a wide range of biotechnology, nanomedicine and bioinformatics applications.\n\nAuthor SummaryThe significance of this work is to expand the applicability of a known metric for describing the function of tiny proteins also called peptides. The Total Similarity Score (TSS) can describe how similar a peptide, or a group of peptides are to another group of sequences with a known or suspected function. A peptide/group of peptides will always have a high TSS if it contains the same or similar amino acids in the same positions. This metric can therefore be used to select peptides for useful functions based purely on conserved amino acids in unknown positions. The greedy search algorithm used to learn how similar amino acids are to each other has been shown to be marginally effective in this larger dataset. Therefore, we argue that the TSS metric is a highly useful one for predicting peptide affinity but a different machine learning algorithm should be applied to make full use of it.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Phage libraries screening on P53: yield improvement by zinc and a new parasites integrating analysis and rationale 95%
- Immuno-informatics Design of a Multimeric Epitope Peptide Based Vaccine Targeting SARS-CoV-2 Spike Glycoprotein 94%
- GraphMHC: neoantigen prediction model applying the graph neural network to molecular structure 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.