Back

Semi-automated retrieval of chemical and phylogenetic information from natural products literature

Coelho, A. C. L.; da Silva, R. R.

2023-06-29 bioinformatics
10.1101/2023.06.28.546864 bioRxiv
Show abstract

Natural products (NPs) are metabolites of great importance due to their fundamental biological role in performing specialized activities, ranging from basic cellular functions to complex ecological interactions. These metabolites have contributed to innovating fields such as agriculture and medicine due to their optimized biological activities, a consequence of evolution. A key factor in ensuring that isolated NPs are novel is to search scientific literature and compare pre-existing chemical entities with the new isolate. Unfortunately, articles are typically not machine-readable, a problem that hinders efficient searching and increases the chances of unintended rediscovery. In addition, the time required to add new compound discoveries to compound databases hinders computational studies on cell metabolism and Quantitative Structure-Activity Relationships (QSAR). Here, we present a modularized tool that uses text mining techniques to retrieve chemical entities and taxonomic mentions present in scientific literature, called NPMINE (Natural Products MINIng). We were able to analyze 55,382 scientific articles from some of the most important applied chemistry journals from Brazil and the world, consistently recovering the expected taxonomic and structural information. This processing resulted in 120,970 unique InChI Keys potentially associated with 21,526 unique species mentioned. Using the PubChem BioAssay database we show how QSAR models can be used to mine active leads. The results indicate that NPMINE not only facilitates natural products cataloging, but also assists in biological source assignment and structure-activity relationships, a time-consuming task, typically performed in low throughput.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Molecules
39 papers in training set
Top 0.1%
12.8%
2
PLOS ONE
5266 papers in training set
Top 16%
12.0%
3
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.1%
9.9%
4
Journal of Cheminformatics
29 papers in training set
Top 0.1%
7.9%
5
Journal of Chemical Information and Modeling
238 papers in training set
Top 1.0%
4.9%
6
Bioinformatics
1204 papers in training set
Top 4%
4.5%
50% of probability mass above
7
BMC Bioinformatics
457 papers in training set
Top 2%
3.5%
8
Scientific Reports
3612 papers in training set
Top 33%
3.2%
9
PLOS Computational Biology
1863 papers in training set
Top 11%
2.8%
10
GigaScience
212 papers in training set
Top 2%
2.4%
11
International Journal of Molecular Sciences
494 papers in training set
Top 7%
1.7%
12
Metabolites
53 papers in training set
Top 0.5%
1.7%
13
Communications Chemistry
48 papers in training set
Top 0.5%
1.7%
14
Database
61 papers in training set
Top 0.5%
1.5%
15
RSC Advances
22 papers in training set
Top 0.4%
1.5%
16
Biomolecules
100 papers in training set
Top 1%
1.3%
17
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.1%
1.3%
18
Briefings in Bioinformatics
354 papers in training set
Top 5%
1.3%
19
Chemical Research in Toxicology
10 papers in training set
Top 0.2%
1.1%
20
Computers in Biology and Medicine
128 papers in training set
Top 3%
1.1%
21
eLife
5828 papers in training set
Top 60%
1.1%
22
Frontiers in Pharmacology
111 papers in training set
Top 2%
1.0%
23
Pharmaceuticals
34 papers in training set
Top 0.9%
1.0%
24
ACS Omega
105 papers in training set
Top 3%
1.0%
25
Nucleic Acids Research
1281 papers in training set
Top 13%
0.9%
26
Scientific Data
209 papers in training set
Top 3%
0.9%
27
Bioinformatics Advances
203 papers in training set
Top 4%
0.9%
28
Genes
144 papers in training set
Top 4%
0.9%
29
Frontiers in Bioinformatics
49 papers in training set
Top 2%
0.6%