Back

EpiESM-GA: Resource-Efficient Protein Foundation Model Features for Equitable B-Cell Epitope Prediction

Gautam, P.; Mitra, P.

2026-06-26 bioinformatics
10.64898/2026.06.22.733745 bioRxiv
Show abstract

Prediction of B-cell epitopes can assist in reducing costly wet-lab screening in vaccine design, diagnostics, and antibody discovery. However, current predictors often suffer from noisy labels, weak generalization, and structure-dependent workflows. Here we present EO_SCPLOWPIC_SCPLOWESM-GA, an efficient sequenceonly pipeline for linear B-cell epitope prediction. Positive and negative peptide examples are collected from IEDB, which provides experimentally tested epitopes and distinguishes positive and negative epitope records based on assay evidence(Vita et al., 2019). Each peptide is encoded with a frozen ESM-2 protein language model: a bidirectional transformer producing amino acid embeddings for downstream structure and function tasks (Lin et al., 2023). Mean-pooled embeddings are further compressed into a compact 420-feature representation with a genetic algorithm and classified with lightweight Random Forest, XGBoost, or MLP heads. This avoids foundation-model fine-tuning, reduces the number of trainable parameters, improves interpretability, and enables low-resource deployment. On an IEDB-derived benchmark, EO_SCPLOWPIC_SCPLOWESM-GA attains 0.880{+/-} 0.004 AUC-ROC, 0.852{+/-} 0.005 PR-AUC, 82.0 {+/-} 0.6% accuracy, 0.79 {+/-} 0.01 F1, and 0.74{+/-} 0.01 MCC, outperforming dense ESM-2 features and baselines LBCE-XGB, EpitopeVec, and BepiPred-2.0 (mean{+/-} std over five independent random seeds). The framework shows how frozen protein foundation models can enable pandemic preparedness, peptide vaccine prioritization, diagnostic antigen screening, and equitable computational immunology.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 0.5%
33.3%
2
Bioinformatics Advances
203 papers in training set
Top 0.2%
10.7%
3
Briefings in Bioinformatics
354 papers in training set
Top 0.9%
7.7%
50% of probability mass above
4
Patterns
78 papers in training set
Top 0.2%
5.3%
5
ImmunoInformatics
12 papers in training set
Top 0.1%
3.9%
6
Nature Machine Intelligence
70 papers in training set
Top 0.7%
3.9%
7
mAbs
32 papers in training set
Top 0.2%
3.3%
8
Nature Communications
5641 papers in training set
Top 38%
2.7%
9
Cell Reports Methods
165 papers in training set
Top 0.9%
2.6%
10
Cell Systems
201 papers in training set
Top 2%
2.1%
11
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
2.1%
12
PLOS Computational Biology
1863 papers in training set
Top 13%
1.9%
13
NAR Genomics and Bioinformatics
242 papers in training set
Top 3%
1.7%
14
Nature Methods
385 papers in training set
Top 4%
1.6%
15
Scientific Reports
3612 papers in training set
Top 67%
1.1%
16
Frontiers in Immunology
638 papers in training set
Top 8%
1.1%
17
eLife
5828 papers in training set
Top 62%
1.0%
18
Cell Genomics
172 papers in training set
Top 4%
1.0%
19
iScience
1154 papers in training set
Top 37%
0.8%
20
BMC Bioinformatics
457 papers in training set
Top 6%
0.8%
21
Journal of Chemical Information and Modeling
238 papers in training set
Top 2%
0.8%
22
GigaScience
212 papers in training set
Top 5%
0.8%
23
Communications Biology
993 papers in training set
Top 32%
0.8%
24
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 46%
0.6%
25
Nature Computational Science
55 papers in training set
Top 2%
0.6%