SynthQA - Hierarchical Machine Learning-based Protein Quality Assessment
Korovnik, M.; Hippe, K.; Hou, J.; Si, D.; Kishaba, K.; Cao, R.
Show abstract
MotivationIt has been a challenge for biologists to determine 3D shapes of proteins from a linear chain of amino acids and understand how proteins carry out lifes tasks. Experimental techniques, such as X-ray crystallography or Nuclear Magnetic Resonance, are time-consuming. This highlights the importance of computational methods for protein structure predictions. In the field of protein structure prediction, ranking the predicted protein decoys and selecting the one closest to the native structure is known as protein model quality assessment (QA), or accuracy estimation problem. Traditional QA methods dont consider different types of features from the protein decoy, lack various features for training machine learning models, and dont consider the relationship between features. In this research, we used multi-scale features from energy score to topology of the protein structure, and proposed a hierarchical architecture for training machine learning models to tackle the QA problem. ResultsWe introduce a new single-model QA method that incorporates multi-scale features from protein structures, utilizes the hierarchical architecture of training machine learning models, and predicts the quality of any protein decoy. Based on our experiment, the new hierarchical architecture is more accurate compared to traditional machine learning-based methods. It also considers the relationship between features and generates additional features so machine learning models can be trained more accurately. We trained our new tool, SynthQA, on the CASP dataset (CASP10 to CASP12), and validated our method on 33 targets from the latest CASP 14 dataset. The result shows that our method is comparable to other state-of-the-art single-model QA methods, and consistently outperforms each of the 14 used features. Availabilityhttps://github.com/Cao-Labs/SynthQA.git Contactcaora@plu.edu
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network 98%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 98%
- Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactions 97%
Similar papers in this journal
- Improved model quality assessment using sequence and structural information by enhanced deep neural networks 97%
- Studying protein-protein interaction through side-chain modeling method OPUS-Mut 97%
- Interpretable and Generalizable Attention-Based Model for Predicting Drug-Target Interaction Using 3D Structure of Protein Binding Sites: SARS-CoV-2 Case Study and in-Lab Validation 96%
Similar papers in this journal
- From Proteins to Ligands: Decoding Deep Learning Methods for Binding Affinity Prediction 95%
- Identification of Family-Specific Features in Cas9 and Cas12 Proteins: A Machine Learning Approach Using Complete Protein Feature Spectrum 95%
- Accurate Conformation Sampling via Protein Structural Diffusion 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.