Back

Sanjeevani: A manually curated anti-cancerous phytochemical database integrated with downstream analysis tools.

Jha, V.; Jha, R.; Shukla, S.; Shingan, S.; Das, G.

2026-06-19 bioinformatics
10.64898/2026.06.15.732344 bioRxiv
Show abstract

BackgroundCancer continues to pose a massive global health burden. While plant-derived phytochemicals offer promising therapeutic leads, existing natural product databases often lack cancer specificity, dataset downloadability, and integrated screening tools. MethodsWe developed Sanjeevani, an integrative web platform cataloguing 4,823 curated anticancer phytochemicals. Using a balanced dataset of 9,646 molecules, we trained Support Vector Machine (SVM), Random Forest, and K-Nearest Neighbours classifiers using a hybrid feature representation of RDKit descriptors and 2048-bit ECFP4 fingerprints. The platform also integrates AutoDock Vina for web-based molecular docking for binding affinity, poses prediction and ADMET-AI for pharmacokinetics estimation. ResultsThe SVM model demonstrated the strongest predictive capability, achieving a top test accuracy of 0.966 and a ROC-AUC of 0.992. Benchmarking across five docking tools confirmed that AutoDock Vina successfully balanced computational automation with literature-consistent binding affinity replication. The final architecture provides rapid interactive 2D/3D visualizations integrated with downstream analysis tools. ConclusionSanjeevani provides an open-access, one-stop pipeline that bridges the gap between raw natural product data and actionable computational screening, accelerating natural product-based oncology drug discovery. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=80 SRC="FIGDIR/small/732344v1_ufig1.gif" ALT="Figure 1"> View larger version (17K): org.highwire.dtl.DTLVardef@77183borg.highwire.dtl.DTLVardef@d7e465org.highwire.dtl.DTLVardef@1d3dfd7org.highwire.dtl.DTLVardef@10cc94d_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.4%
13.2%
2
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.1%
13.2%
3
Bioinformatics
1204 papers in training set
Top 3%
8.2%
4
Scientific Reports
3612 papers in training set
Top 8%
7.6%
5
PLOS ONE
5266 papers in training set
Top 32%
4.6%
6
Journal of Cheminformatics
29 papers in training set
Top 0.2%
4.2%
50% of probability mass above
7
BMC Bioinformatics
457 papers in training set
Top 2%
4.2%
8
Briefings in Bioinformatics
354 papers in training set
Top 2%
3.6%
9
PLOS Computational Biology
1863 papers in training set
Top 9%
3.4%
10
Molecules
39 papers in training set
Top 0.5%
2.0%
11
Database
61 papers in training set
Top 0.4%
2.0%
12
GigaScience
212 papers in training set
Top 2%
1.8%
13
Frontiers in Pharmacology
111 papers in training set
Top 1%
1.8%
14
ACS Omega
105 papers in training set
Top 1%
1.8%
15
Journal of Translational Medicine
57 papers in training set
Top 0.9%
1.6%
16
Computers in Biology and Medicine
128 papers in training set
Top 3%
1.2%
17
Bioinformatics Advances
203 papers in training set
Top 4%
1.2%
18
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.2%
1.1%
19
Communications Chemistry
48 papers in training set
Top 1%
1.0%
20
Scientific Data
209 papers in training set
Top 3%
0.9%
21
Frontiers in Oncology
103 papers in training set
Top 3%
0.6%
22
Journal of Medicinal Chemistry
77 papers in training set
Top 0.9%
0.6%
23
Pharmaceuticals
34 papers in training set
Top 2%
0.5%
24
Journal of Biomolecular Structure and Dynamics
43 papers in training set
Top 1%
0.5%
25
BioData Mining
22 papers in training set
Top 1%
0.5%
26
Plant Communications
36 papers in training set
Top 1%
0.5%
27
International Journal of Molecular Sciences
494 papers in training set
Top 18%
0.5%