Back

HALPred-B: Host-Aware Linear B-Cell Epitope Prediction: Challenges, Limitations, and Variability Across Species

Gautam, P.; Mitra, P.; Sinha, I.

2026-06-26 bioinformatics
10.64898/2026.06.22.733770 bioRxiv
Show abstract

Predicting linear B-cell epitopes is a basic immunoinformatics task that has a direct impact on vaccine design and antibody engineering. Recent advances in machine learning have improved predictive performance, but most existing approaches are trained on aggregated datasets and assume that antigenic patterns are conserved across host organisms. This assumption ignores the immunological variability depending on the host and prevents generalizing the model across species. This is the first systematic host-wise evaluation where we present a systematic machine learning-based analysis of host-aware linear B-cell epitope prediction using curated datasets from the Immune Epitope Database (IEDB). We build separate datasets for human, mouse, and non-human primate hosts and assess several classification models, including Random Forest, Support Vector Machine (SVM), Gradient Boosting, XGBoost, and K-Nearest Neighbors (KNN). The models exploit feature representations derived from sequences, such as AAIndex descriptors, biochemical properties from ExPASy, and dipeptide composition. Our results show that predictive performance differs substantially across hosts. Models achieve up to 86.07% accuracy and 0.93 ROC-AUC on human datasets but lower performance on mouse and non-human primate datasets. This gap underlies dataset bias and sequence distribution differences, as well as the inability of existing features to capture host-specific immunological context. These results indicate that the prediction of linear B-cell epitopes is intrinsically host-specific, and a single global model does not generalize well across species. We propose to incorporate host-aware modeling strategies and organism-specific features for enhanced predictive reliability and biological relevance.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
ImmunoInformatics
12 papers in training set
Top 0.1%
28.6%
2
mAbs
32 papers in training set
Top 0.1%
8.9%
3
Frontiers in Immunology
638 papers in training set
Top 2%
7.9%
4
PLOS Computational Biology
1863 papers in training set
Top 4%
7.9%
50% of probability mass above
5
Scientific Reports
3612 papers in training set
Top 18%
5.2%
6
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.3%
7
PLOS ONE
5266 papers in training set
Top 35%
3.5%
8
Bioinformatics Advances
203 papers in training set
Top 2%
3.2%
9
Bioinformatics
1204 papers in training set
Top 5%
3.2%
10
Nature Communications
5641 papers in training set
Top 38%
2.6%
11
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
2.4%
12
iScience
1154 papers in training set
Top 13%
2.1%
13
NAR Genomics and Bioinformatics
242 papers in training set
Top 3%
1.3%
14
GigaScience
212 papers in training set
Top 3%
1.1%
15
Cell Systems
201 papers in training set
Top 4%
1.1%
16
Frontiers in Bioinformatics
49 papers in training set
Top 1%
1.0%
17
eLife
5828 papers in training set
Top 63%
0.9%
18
Protein Science
246 papers in training set
Top 3%
0.8%
19
BMC Bioinformatics
457 papers in training set
Top 6%
0.8%
20
BMC Genomics
406 papers in training set
Top 9%
0.6%
21
Science China Life Sciences
29 papers in training set
Top 0.7%
0.6%
22
Communications Biology
993 papers in training set
Top 35%
0.6%
23
Patterns
78 papers in training set
Top 3%
0.6%
24
Cell Genomics
172 papers in training set
Top 4%
0.6%