Back

Bioinf-Farma: supervised integration of epitope prediction and recombinant protein developability for automated vaccine candidate prioritization

Bondi, H.; Crespi, M.; Orlando, M.; Lescai, F.; Serapian, S. A.; Colombo, G.; Fasano, M.; Pollegioni, L.; Molla, G.

2026-06-18 bioinformatics
10.64898/2026.06.15.732271 bioRxiv
Show abstract

Vaccine antigen discovery requires prioritizing protein candidates according to both immunogenic potential and recombinant expression feasibility. These properties are typically evaluated using separate computational tools, requiring researchers to integrate heterogeneous outputs through ad hoc workflows. Here, we present BIOINF-farma, a modular platform integrating epitope prediction and developability assessment for rational antigen selection within a unified environment. Candidates can be submitted as amino acid sequences or three-dimensional structures. When experimental structures are unavailable, BIOINF-farma automatically searches for models in AlphaFold DB or performs structure prediction using Boltz-2, ensuring a standardized structural representation for downstream analyses. Antigenicity is quantified by combining structure-based conformational epitope signals (MLCE/REBELOT-BEPPE) and sequence-based linear epitope propensity scores (BepiPred 3.0) into a protein-level Antigenicity Score, with a classification threshold optimized on a manually curated validation dataset. Developability is evaluated through two supervised Random Forest meta-learners that integrate three solubility predictors (DeepSoluE, SoluProt, Protein-Sol) and three thermal stability predictors (TemStaPro, ProLaTherm, BertThermo), whose outputs are combined into an Expression Efficiency Score (EES). By integrating complementary predictive signals, the meta-learning framework achieves greater accuracy and robustness than individual predictors while maintaining performance across a broad range of sequence identities. The Antigenicity Score effectively discriminates antigenic from non-antigenic proteins with a large effect size, whereas EES successfully distinguishes soluble from insoluble outcomes on an independent panel of recombinant proteins expressed in Escherichia coli. BIOINF-farma jointly assesses antigenicity and expression feasibility within a single framework. Its modular architecture facilitates the incorporation of future predictive methods, while its web-based interface makes the full pipeline accessible to users without programming expertise, supporting rapid candidate triage in vaccine research and emerging pathogen responses. Author SummaryVaccine development begins with a critical step: identifying, among the many proteins encoded in a pathogen genome, those most suitable as candidate antigens. A promising candidate must satisfy two requirements that are rarely evaluated together. It must be recognized by the immune system, so that vaccination elicits a protective response; and it must be amenable to recombinant production, since antigens that cannot be obtained in sufficient quantity and quality are of limited practical use. Current computational tools typically address only one of these aspects, and researchers must integrate their outputs manually, through procedures that are time-consuming and prone to inconsistency. We developed BIOINF-farma, an automated platform that brings these two assessments into a single analytical framework. Starting from a protein sequence or an experimental structure, the platform retrieves or predicts a three-dimensional model, evaluates the proteins antigenic potential by combining complementary epitope predictors, and estimates its expression feasibility by integrating multiple solubility and stability predictors through supervised machine learning. A web-based interface makes the full workflow available to experimental immunologists and vaccine developers without requiring computational expertise, supporting rational candidate prioritization in routine vaccine research and during emerging pathogen responses.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
mAbs
32 papers in training set
Top 0.1%
10.8%
2
Briefings in Bioinformatics
354 papers in training set
Top 0.6%
9.6%
3
ImmunoInformatics
12 papers in training set
Top 0.1%
8.8%
4
Bioinformatics
1204 papers in training set
Top 3%
8.8%
5
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.4%
6.2%
6
PLOS Computational Biology
1863 papers in training set
Top 7%
5.4%
7
Bioinformatics Advances
203 papers in training set
Top 1.0%
5.4%
50% of probability mass above
8
Frontiers in Immunology
638 papers in training set
Top 3%
4.3%
9
BMC Bioinformatics
457 papers in training set
Top 2%
4.0%
10
PLOS ONE
5266 papers in training set
Top 35%
4.0%
11
Nature Communications
5641 papers in training set
Top 36%
3.2%
12
Scientific Reports
3612 papers in training set
Top 35%
3.2%
13
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
2.4%
14
Protein Science
246 papers in training set
Top 2%
2.3%
15
Cell Reports Methods
165 papers in training set
Top 1%
2.0%
16
GigaScience
212 papers in training set
Top 2%
1.9%
17
Journal of Molecular Biology
232 papers in training set
Top 2%
1.7%
18
Frontiers in Bioinformatics
49 papers in training set
Top 1%
0.9%
19
Nucleic Acids Research
1281 papers in training set
Top 14%
0.8%
20
Cell Systems
201 papers in training set
Top 5%
0.8%
21
BMC Genomics
406 papers in training set
Top 10%
0.6%
22
International Journal of Molecular Sciences
494 papers in training set
Top 18%
0.6%
23
Communications Biology
993 papers in training set
Top 36%
0.6%