Detection of native and mirror protein structures based on Ramachandran plot analysis by interpretable machine learning models
Villmann, T.; Abel, J.; Bohnsack, K. S.; Kaden, M.; Weber, M.; Leberecht, C.
Show abstract
In this contribution the discrimination between native and mirror models of proteins according to their chirality is tackled based on the structural protein information. This information is contained in the Ramachandran plots of the protein models. We provide an approach to classify those plots by means of an interpretable machine learning classifier - the Generalized Matrix Learning Vector Quantizer. Applying this tool, we are able to distinguish with high accuracy between mirror and native structures just evaluating the Ramachandran plots. The classifier model provides additional information regarding the importance of regions, e.g. -helices and {beta}-strands, to discriminate the structures precisely. This importance weighting differs for several considered protein classes.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SARS-CoV-2 protein structure and sequence mutations: evolutionary analysis and effects on virus variants SARS-CoV-2 protein structure and sequence mutations: 95%
- Statistical potentials from the Gaussian scaling behaviour of chain fragments buried within protein globules 94%
- Classification of protein binding ligands using structural dispersion of binding site atoms from principal axes 94%
Similar papers in this journal
- Predicting Affinity Through Homology (PATH): Interpretable Binding Affinity Prediction with Persistent Homology 95%
- Novel, provable algorithms for efficient ensemble-based computational protein design and their application to the redesign of the c-Raf-RBD:KRas protein-protein interface 95%
- Interpretable Pairwise Distillations for Generative Protein Sequence Models 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.