Supervised learning of protein thermal stability using sequence mining and distribution statistics of network centrality
Sharma, A.; Bagler, G.; Bera, D.
Show abstract
MotivationIt is expected that the difference in the thermal stability of mesophilic and thermophilic proteins arises, in part at least, from the differences in their molecular structures and amino acid compositions. Existing machine learning approaches for supervised classification of proteins rely on the features derived from the structural networks and the amino acid sequences. However, the network features used leave out several important network centrality values, the statistic used is a simple average and the sequence features used are hand-picked leading to an accuracy of 90%.\n\nResultsWe show that discriminating sub-sequences of the amino acid sequences can significantly improve classification accuracy compared to the existing approaches of counting amino acids, di-peptide or even tri-peptide bonds. We identify notions of network centrality, specifically that depends on the distances between C atoms, that appears to correlate better with thermal stability compared to the existing network features. We also show how to generate better statistics from the node- and edge-wise centrality values that more accurately captures the variations in their values for different types of proteins. These improved feature selection techniques make it possible to classify between thermophilic and mesophilic proteins with 96% accuracy and 99% area under ROC.\n\nAvailabilityThe dataset and source code used are available at https://github.com/ankits0207/Protein_Classification_BIO699\n\nContactdbera@iiitd.ac.in\n\nonline.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Revisiting structural organization of proteins at high temperature from network perspective 95%
- SubFeat: Feature Subspacing Ensemble Classifier for Function Prediction of DNA, RNA and Protein Sequences 94%
- ProS-GNN: Predicting effects of mutations on protein stability using graph neural networks 93%
Similar papers in this journal
Similar papers in this journal
- Designing of thermostable proteins with a desired melting temperature 96%
- Predicting interchain contacts for homodimeric and homomultimeric protein complexes using multiple sequence alignments of monomers and deep learning 95%
- Influence of spatial structure on protein damage susceptibility - A bioinformatics approach 95%
Similar papers in this journal
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 93%
- Towards a comprehensive view of the pocketome universe - biological implications and algorithmic challenges. 93%
- A Novel Riboswitch Classification based on Imbalanced Sequences achieved by Machine Learning 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.