Back

Detection of native and mirror protein structures based on Ramachandran plot analysis by interpretable machine learning models

Villmann, T.; Abel, J.; Bohnsack, K. S.; Kaden, M.; Weber, M.; Leberecht, C.

2020-09-03 bioinformatics
10.1101/2020.09.03.280701 bioRxiv
Show abstract

In this contribution the discrimination between native and mirror models of proteins according to their chirality is tackled based on the structural protein information. This information is contained in the Ramachandran plots of the protein models. We provide an approach to classify those plots by means of an interpretable machine learning classifier - the Generalized Matrix Learning Vector Quantizer. Applying this tool, we are able to distinguish with high accuracy between mirror and native structures just evaluating the Ramachandran plots. The classifier model provides additional information regarding the importance of regions, e.g. -helices and {beta}-strands, to discriminate the structures precisely. This importance weighting differs for several considered protein classes.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.