Back

Development of Deep-Learning Models that Predict Quantitative Protein-Ligand Interac-tions in Glycobiology as a part of a Capstone Course

Yin, H.; Liu, W.; Zhou, W.; Chang, Z.; Carpenter, E. J.; Satyajith, A.; Haregu, S.; Greiner, R.; Derda, R.

2026-06-24 bioinformatics
10.64898/2026.06.19.733466 bioRxiv
Show abstract

Glycans coat the surface of all cells, and every glycan is recognised by specific glycan-binding pro-teins (GBPs). There are no general tools that can accurately estimate the binding strength between glycan and GBP from the amino acid sequence of the GBP and the molecular structure of the glycan, represented as SMILES string. We describe models for predicting such binding strengths developed as a part of a Capstone Course at the University of Alberta. The models are trained on a dataset that combines BindingDB, a published database of small-molecule protein interactions, and data from glycan arrays measured by Consortium of Functional Glycomics (CFG). In this hybrid dataset of protein-ligand interactions the ligands are both glycans from CFG and small molecules from BindingDB; similarly, proteins include GBP and proteins from BindingDB. Three models are presented (i) ProMax which fuses ESM-2, MolFormer, and MolCLR features; (ii) APEX which constrains learning to a predetermined form, a physical model of binding; (iii) UltraMax adds inter-atomic distances for the ligands. To address the dataset's severe long-tail distribution, the models employ tail-aware losses for rare high-binding instances. Trained and evaluated on approximately one million protein--ligand pairs using hold-out splits for unseen molecules, the three models provide a unified framework for quantitative glycan-protein binding prediction. We observed that learning glycan-protein binding is harder than the similar task of learning small-molecule-protein interactions. Simple mirror-inversion tests led us to postulate that insufficient use of chiral features is an important source of difficulty in learning these interactions.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 1%
19.0%
2
PLOS Computational Biology
1863 papers in training set
Top 2%
13.3%
3
Glycobiology
35 papers in training set
Top 0.1%
8.1%
4
Bioinformatics Advances
203 papers in training set
Top 0.5%
6.9%
5
Nature Methods
385 papers in training set
Top 3%
3.5%
50% of probability mass above
6
Nature Machine Intelligence
70 papers in training set
Top 0.8%
3.3%
7
Molecular & Cellular Proteomics
25 papers in training set
Top 0.1%
3.3%
8
Frontiers in Bioinformatics
49 papers in training set
Top 0.1%
3.3%
9
Patterns
78 papers in training set
Top 0.7%
2.8%
10
Scientific Reports
3612 papers in training set
Top 42%
2.5%
11
Nature Communications
5641 papers in training set
Top 42%
2.1%
12
mAbs
32 papers in training set
Top 0.2%
2.0%
13
Briefings in Bioinformatics
354 papers in training set
Top 4%
2.0%
14
Cell Systems
201 papers in training set
Top 2%
1.8%
15
Journal of Chemical Information and Modeling
238 papers in training set
Top 2%
1.7%
16
PLOS ONE
5266 papers in training set
Top 54%
1.2%
17
BMC Bioinformatics
457 papers in training set
Top 5%
1.2%
18
eLife
5828 papers in training set
Top 56%
1.2%
19
ACS Omega
105 papers in training set
Top 3%
1.0%
20
Computational and Structural Biotechnology Journal
242 papers in training set
Top 6%
1.0%
21
Journal of Computational Biology
48 papers in training set
Top 1%
0.9%
22
ImmunoInformatics
12 papers in training set
Top 0.2%
0.6%
23
Nucleic Acids Research
1281 papers in training set
Top 14%
0.6%
24
Cell Reports
1498 papers in training set
Top 28%
0.6%
25
Cell Reports Methods
165 papers in training set
Top 4%
0.6%
26
BMC Genomics
406 papers in training set
Top 9%
0.6%