Back

Supervised learning of protein thermal stability using sequence mining and distribution statistics of network centrality

Sharma, A.; Bagler, G.; Bera, D.

2019-09-24 bioinformatics
10.1101/777177 bioRxiv
Show abstract

MotivationIt is expected that the difference in the thermal stability of mesophilic and thermophilic proteins arises, in part at least, from the differences in their molecular structures and amino acid compositions. Existing machine learning approaches for supervised classification of proteins rely on the features derived from the structural networks and the amino acid sequences. However, the network features used leave out several important network centrality values, the statistic used is a simple average and the sequence features used are hand-picked leading to an accuracy of 90%.\n\nResultsWe show that discriminating sub-sequences of the amino acid sequences can significantly improve classification accuracy compared to the existing approaches of counting amino acids, di-peptide or even tri-peptide bonds. We identify notions of network centrality, specifically that depends on the distances between C atoms, that appears to correlate better with thermal stability compared to the existing network features. We also show how to generate better statistics from the node- and edge-wise centrality values that more accurately captures the variations in their values for different types of proteins. These improved feature selection techniques make it possible to classify between thermophilic and mesophilic proteins with 96% accuracy and 99% area under ROC.\n\nAvailabilityThe dataset and source code used are available at https://github.com/ankits0207/Protein_Classification_BIO699\n\nContactdbera@iiitd.ac.in\n\nonline.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

1
Computational Biology and Chemistry
28 papers in training set
Top 0.1%
15.2%
2
Bioinformatics
1204 papers in training set
Top 3%
7.9%
3
Scientific Reports
3612 papers in training set
Top 12%
6.3%
4
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.7%
4.9%
5
PLOS Computational Biology
1863 papers in training set
Top 8%
4.4%
6
PeerJ
308 papers in training set
Top 2%
3.3%
7
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.3%
8
Frontiers in Bioinformatics
49 papers in training set
Top 0.1%
3.3%
9
Computers in Biology and Medicine
128 papers in training set
Top 1%
3.3%
50% of probability mass above
10
PLOS ONE
5266 papers in training set
Top 37%
3.3%
11
BMC Bioinformatics
457 papers in training set
Top 3%
2.7%
12
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 0.5%
2.4%
13
International Journal of Biological Macromolecules
76 papers in training set
Top 0.7%
2.1%
14
Biosystems
31 papers in training set
Top 0.2%
2.0%
15
ACS Omega
105 papers in training set
Top 1%
1.7%
16
Journal of Chemical Information and Modeling
238 papers in training set
Top 2%
1.7%
17
Journal of Computational Biology
48 papers in training set
Top 0.6%
1.7%
18
Bioinformatics Advances
203 papers in training set
Top 3%
1.7%
19
Frontiers in Molecular Biosciences
102 papers in training set
Top 1%
1.1%
20
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 0.8%
1.1%
21
Journal of Computational Chemistry
13 papers in training set
Top 0.2%
1.1%
22
Physica A: Statistical Mechanics and its Applications
13 papers in training set
Top 0.2%
1.1%
23
Physical Biology
46 papers in training set
Top 0.6%
1.1%
24
Journal of Theoretical Biology
162 papers in training set
Top 2%
1.1%
25
Biomolecules
100 papers in training set
Top 2%
1.1%
26
Frontiers in Genetics
230 papers in training set
Top 4%
1.1%
27
Molecules
39 papers in training set
Top 1%
1.0%
28
Biology Methods and Protocols
61 papers in training set
Top 2%
1.0%
29
International Journal of Molecular Sciences
494 papers in training set
Top 13%
1.0%
30
Protein Science
246 papers in training set
Top 3%
0.9%