Back

THEOBROMA: an aggregated open database of 1.13 million natural products with per-compound license auditing, three-tier classification, and stereochemistry-aware deduplication

Klamt, T.; Jaczkowski, A.; Franke, J.; Nejdl, W.

2026-06-16 bioinformatics
10.64898/2026.06.12.731585 bioRxiv
Show abstract

Natural products remain one of the most productive sources of pharmacologically active compounds for drug discovery, yet the current open aggregator landscape attributes licenses at database rather than compound granularity, with consequences that have become tangible as the field grows. A recent relicensing event in one constituent source (the September 2024 transition of the Natural Products Atlas to CC BY-NC 4.0) demonstrates how database-level licensing propagates across an aggregate and motivates the per-compound audit framework presented here. The same peer cohort separately leaves classification provenance and stereoisomer-family relations coarser than either layer warrants. THEOBROMA, accessible at https://theobroma.l3s.uni-hannover.de, integrates 1,133,004 natural products from 29 open sources under a per-compound license audit that resolves each compounds license tier across all attesting sources under a most-restrictive-wins rule, identifying 900,170 compounds (79.4%) under open-use licenses and exposing the per-source attestation chain and resolved tier through a dedicated audit endpoint and a query-time license filter. A three-tier classification stratifies 89.3% coverage into 35.1% curated, 43.9% high-confidence inferred, and 10.3% exploratory tiers, with 486,215 stereoisomer families preserved by full 27-character InChIKey deduplication and exposed via a dedicated /api/stereoisomers/<comp_id> endpoint and a radial-family display. Per-compound license provenance is the primary differentiator. Classification stratification and stereoisomer-family exposure add finer-grained access to two related axes, supporting license-compatible virtual screening and isomer-specific bioactivity analysis at corpus scale. As an evolving open resource, THEOBROMA pairs continuous pipeline maintenance with interactive geographic, taxonomic, and chemical-space exploration.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 17%
10.7%
2
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.5%
9.8%
3
Communications Chemistry
48 papers in training set
Top 0.1%
8.0%
4
Nature Methods
385 papers in training set
Top 1%
6.8%
5
Nature
645 papers in training set
Top 3%
5.6%
6
Nucleic Acids Research
1281 papers in training set
Top 4%
4.4%
7
Nature Biotechnology
172 papers in training set
Top 0.9%
4.4%
8
Bioinformatics
1204 papers in training set
Top 5%
3.3%
50% of probability mass above
9
Journal of Cheminformatics
29 papers in training set
Top 0.3%
2.8%
10
PLOS ONE
5266 papers in training set
Top 48%
1.7%
11
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 27%
1.7%
12
Nature Protocols
33 papers in training set
Top 0.2%
1.7%
13
Scientific Data
209 papers in training set
Top 2%
1.5%
14
Science
477 papers in training set
Top 6%
1.4%
15
Chemical Science
73 papers in training set
Top 1%
1.3%
16
Briefings in Bioinformatics
354 papers in training set
Top 5%
1.3%
17
Cell Systems
201 papers in training set
Top 3%
1.1%
18
Cell Genomics
172 papers in training set
Top 3%
1.1%
19
Computational and Structural Biotechnology Journal
242 papers in training set
Top 5%
1.1%
20
Advanced Science
286 papers in training set
Top 7%
1.1%
21
PLOS Computational Biology
1863 papers in training set
Top 17%
1.1%
22
Nature Chemical Biology
119 papers in training set
Top 2%
1.1%
23
Database
61 papers in training set
Top 0.6%
1.1%
24
Communications Biology
993 papers in training set
Top 21%
1.1%
25
iScience
1154 papers in training set
Top 25%
1.1%
26
eLife
5828 papers in training set
Top 59%
1.1%
27
Patterns
78 papers in training set
Top 2%
1.0%
28
Cell Chemical Biology
94 papers in training set
Top 1%
1.0%
29
Scientific Reports
3612 papers in training set
Top 73%
0.9%
30
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.3%
0.9%