Back

An Open Benchmark for Systems Vaccinology: Insights from the CMI-PB Challenges

Shinde, P.; Willemsen, L.; Lee, J.; Orfield, S.; Ren, Z.; Aoki, M.; Thrupp, N.; Gupta, A.; Wu, C.-C.; Mao, L.; Li, C.; Tan, Y.; Nguyen, T. A.; Chang, N.-S.; Schafer, P. S. L.; Xing, J.; Can Ali Marandi, C.; Sabuwala, B.; Reyna, J.; Gygi, J. P.; Ha, B.; Overton, J. A.; Einav, T.; Greenbaum, J. A.; Guan, L.; Kojima, M.; Ay, F.; Grant, B.; Kleinstein, S. H.; Peters, B.

2026-08-28 immunology
10.64898/2026.08.25.746820 bioRxiv
Show abstract

Systems vaccinology approaches have identified factors affecting vaccine responses in multiple studies, but the ability of computational models to generalize these findings to unseen data remains unclear. We established a community resource to create and compare models predicting B. pertussis booster vaccination responses and put such modeling approaches to the test. We compiled multi-modal experimental training data from three independent cohorts (n=117 individuals), and asked investigators to predict vaccine responses in a cohort of 54 newly recruited individuals using only their pre-booster vaccination data. We benchmarked a total of 107 computational models. Top-performing models were characterized by workflows that prioritized rigorous data preprocessing, robust imputation of missing data, and the use of multi-omics integration or non-linear machine learning. We identified pre-existing antigen-specific antibody titers and baseline monocyte frequencies as the most consistent predictors of post-vaccination immunity, highlighting the dominant role of individual immune setpoints. We established the resulting datasets and evaluation framework as a community resource to advance predictive immunology and facilitate personalized vaccination strategies.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.