Back

Bioactivity assessment of natural compounds using machine learning models based on drug target similarity

Periwal, V.; Bassler, S.; Andrejev, S.; Gabrielli, N.; Typas, A.; Patil, K. R.

2020-11-08 bioinformatics
10.1101/2020.11.06.371112 bioRxiv
Show abstract

Natural compounds constitute a rich resource of potential small-molecule therapeutics. While experimental access to this resource is limited due to its vast diversity and difficulties in systematic purification, computational assessment of structural similarity with known therapeutic molecules offers a scalable approach. Here, we assessed functional similarity between natural compounds and approved drugs by combining multiple chemical similarity metrics and physicochemical properties through a random forest model. As a training set, we used pair-wise similarity between 1410 drugs in terms of their shared protein targets. The resulting model featured high performance metrics (matthews correlation coefficient of 0.81, and balanced accuracy of 0.91) suggesting that it well-captured the structure-activity relation. The model was then used to predict protein targets of circa 11k natural compounds by comparing them with the drugs. This revealed therapeutic potential of several natural compounds, including those with support from previously published sources as well as those hitherto unexplored. We experimentally validated one of the predicted links activities, viz., Cox-1 inhibition by 5-methoxysalicylic acid, a molecule commonly found in tea, herbs and spices. In contrast, another natural compound, 4-isopropylbenzoic acid, which showed a higher similarity when considering the most weighted similarity metric but was not picked by the random forest model, did not inhibit Cox-1. Our results demonstrate the utility of a machine-learning approach combining multiple chemical features for uncovering protein binding potential of natural compounds.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Journal of Cheminformatics
29 papers in training set
Top 0.1%
18.4%
2
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.4%
12.6%
3
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.2%
7.8%
4
Scientific Reports
3612 papers in training set
Top 16%
5.5%
5
Bioinformatics
1204 papers in training set
Top 4%
4.8%
6
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.0%
50% of probability mass above
7
iScience
1154 papers in training set
Top 4%
4.0%
8
PLOS Computational Biology
1863 papers in training set
Top 10%
3.2%
9
Frontiers in Pharmacology
111 papers in training set
Top 0.9%
2.8%
10
PLOS ONE
5266 papers in training set
Top 42%
2.4%
11
Communications Chemistry
48 papers in training set
Top 0.3%
2.4%
12
Frontiers in Chemistry
16 papers in training set
Top 0.1%
1.9%
13
Molecules
39 papers in training set
Top 0.6%
1.9%
14
ACS Omega
105 papers in training set
Top 1%
1.9%
15
Bioinformatics Advances
203 papers in training set
Top 3%
1.7%
16
Computers in Biology and Medicine
128 papers in training set
Top 2%
1.7%
17
eLife
5828 papers in training set
Top 53%
1.4%
18
International Journal of Molecular Sciences
494 papers in training set
Top 11%
1.1%
19
Pharmaceuticals
34 papers in training set
Top 0.9%
1.1%
20
ACS Pharmacology & Translational Science
40 papers in training set
Top 0.7%
0.9%
21
European Journal of Pharmacology
15 papers in training set
Top 0.6%
0.8%
22
Biomolecules
100 papers in training set
Top 3%
0.8%
23
Database
61 papers in training set
Top 1%
0.6%
24
Journal of Medicinal Chemistry
77 papers in training set
Top 1%
0.6%
25
The Pharmacogenomics Journal
11 papers in training set
Top 0.3%
0.6%
26
BMC Bioinformatics
457 papers in training set
Top 6%
0.6%