Back

Interpretable sequence-based machine learning consolidates candidate H3N2 hemagglutinin antigenic sites

Meyer, A. G.; Santillana, M.

2026-05-01 bioinformatics
10.64898/2026.04.28.721429 bioRxiv
Show abstract

Vaccine strain selection for seasonal influenza A(H3N2) depends on knowing which hemagglutinin (HA) substitutions are most likely to erode neutralizing antibody recognition, yet published antigenic site sets disagree substantially on which positions matter most. We applied interpretable gradient-boosted tree models with SHAP-based site attribution to two complementary hemagglutination inhibition (HI) datasets to produce a more consolidated ranking of candidate antigenic positions. Models trained on a Neher/Bedford benchmark dataset recover the canonical cluster-transition sites established by prior analyses. Moreover, after filtering the WIC dataset for confounding factors, our models recover the majority of positions from four major prior reference sets (Koel, Neher/Bedford, Harvey, and Shah) and improve concordance between rankings derived from the Neher/Bedford and WIC datasets. Rankings from our models also agree more strongly with models trained to predict sampling time or passage identity than with standard evolutionary metrics used to detect diversifying selection. Our results show that interpretable sequence-based models can provide a more integrative ranking of candidate antigenic positions across different data sources and modeling approaches. This work should aid efforts to prioritize H3N2 substitutions for epidemic surveillance. Significance StatementEvery year, health authorities must update the seasonal flu vaccine to account for mutations in influenza A(H3N2) that allow the virus to escape existing immunity. Knowing which specific positions in the hemagglutinin protein drive this immune escape is essential for evaluating newly emerging variants, but published studies disagree substantially on which positions matter most. We show that interpretable machine learning models applied to two hemagglutination inhibition datasets, the Neher/Bedford benchmark dataset and the larger WHO Collaborating Centre dataset, can help to resolve the disagreements. The models recover canonical cluster-transition sites from the Neher/Bedford benchmark data, and show that our analysis approach with the WIC data improves concordance across several prior rankings produced from distinct datasets and modeling approaches. The resulting rankings provide a practical, consolidated reference for prioritizing hemagglutinin mutations most likely to affect vaccine effectiveness.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.