Back

SynthQA - Hierarchical Machine Learning-based Protein Quality Assessment

Korovnik, M.; Hippe, K.; Hou, J.; Si, D.; Kishaba, K.; Cao, R.

2021-01-30 bioinformatics
10.1101/2021.01.28.428710 bioRxiv
Show abstract

MotivationIt has been a challenge for biologists to determine 3D shapes of proteins from a linear chain of amino acids and understand how proteins carry out lifes tasks. Experimental techniques, such as X-ray crystallography or Nuclear Magnetic Resonance, are time-consuming. This highlights the importance of computational methods for protein structure predictions. In the field of protein structure prediction, ranking the predicted protein decoys and selecting the one closest to the native structure is known as protein model quality assessment (QA), or accuracy estimation problem. Traditional QA methods dont consider different types of features from the protein decoy, lack various features for training machine learning models, and dont consider the relationship between features. In this research, we used multi-scale features from energy score to topology of the protein structure, and proposed a hierarchical architecture for training machine learning models to tackle the QA problem. ResultsWe introduce a new single-model QA method that incorporates multi-scale features from protein structures, utilizes the hierarchical architecture of training machine learning models, and predicts the quality of any protein decoy. Based on our experiment, the new hierarchical architecture is more accurate compared to traditional machine learning-based methods. It also considers the relationship between features and generates additional features so machine learning models can be trained more accurately. We trained our new tool, SynthQA, on the CASP dataset (CASP10 to CASP12), and validated our method on 33 targets from the latest CASP 14 dataset. The result shows that our method is comparable to other state-of-the-art single-model QA methods, and consistently outperforms each of the 14 used features. Availabilityhttps://github.com/Cao-Labs/SynthQA.git Contactcaora@plu.edu

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.