Back

A control-validated pan-proteome deep-learning pipeline nominates GPR35 as a candidate target of the orphan bacterial metabolite ligiamycin A

Martin, J.

2026-07-06 bioinformatics
10.64898/2026.07.01.735807 bioRxiv
Show abstract

Most microbial natural products with documented bioactivity lack an identified molecular target, which limits their development. We present an open, control-validated computational pipeline for natural-product target hypothesis generation. It combines a pan-proteome deep-learning drug-target interaction (DTI) model (a graph neural-network ligand encoder, an ESM-2 protein language-model encoder, and bidirectional cross-attention) with bias-corrected ranking and control-anchored molecular docking. Applying it to ligiamycin A, a 2022-described Streptomyces/Achromobacter co-culture decalin-amino-maleimide with no reported target, we find that the predicted interactions of the compound are dominated by class-A G-protein-coupled receptors. Using a drug with a known target (losartan) we identify and correct a frequent-hitter bias in the raw model; after correction the standout candidates are uniformly class-A GPCRs, led by the orphan receptor GPR35. Structure-based docking with matched positive and negative controls across three candidates corroborates GPR35 specifically: ligiamycin A scores comparably to the known GPR35 agonist zaprinast at the agonist pocket (-8.1 vs -8.3 kcal/mol; non-binder floor -5.5), whereas FFAR1 is excluded and histamine H2 is inconclusive. We propose GPR35 as a prioritized, experimentally testable target and release the workflow as a reusable tool. The result is a computational hypothesis that requires experimental validation.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.1%
41.1%
2
Journal of Cheminformatics
29 papers in training set
Top 0.1%
8.2%
3
Communications Chemistry
48 papers in training set
Top 0.1%
7.0%
50% of probability mass above
4
Scientific Reports
3612 papers in training set
Top 14%
5.7%
5
Nature Communications
5641 papers in training set
Top 37%
2.9%
6
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
2.5%
7
Briefings in Bioinformatics
354 papers in training set
Top 4%
2.2%
8
PLOS ONE
5266 papers in training set
Top 44%
2.2%
9
PLOS Computational Biology
1863 papers in training set
Top 12%
2.2%
10
Bioinformatics
1204 papers in training set
Top 7%
1.8%
11
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.1%
1.8%
12
Chemical Science
73 papers in training set
Top 1%
1.6%
13
Molecules
39 papers in training set
Top 0.9%
1.2%
14
Bioinformatics Advances
203 papers in training set
Top 4%
1.2%
15
iScience
1154 papers in training set
Top 33%
0.9%
16
eLife
5828 papers in training set
Top 63%
0.9%
17
Communications Biology
993 papers in training set
Top 28%
0.9%
18
Advanced Science
286 papers in training set
Top 10%
0.6%
19
ACS Omega
105 papers in training set
Top 4%
0.6%
20
International Journal of Molecular Sciences
494 papers in training set
Top 16%
0.6%
21
NAR Genomics and Bioinformatics
242 papers in training set
Top 5%
0.6%
22
mAbs
32 papers in training set
Top 0.6%
0.5%
23
Frontiers in Pharmacology
111 papers in training set
Top 4%
0.5%
24
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 46%
0.5%