Back

Fine-tuned Protein Language Model Identifies Antigen-specific B Cell Receptors from Immune Repertoires

Paco, K.; Mendivil, M. P.; Zhang, Z.; Zebardast, S.; Davila, C.; Mooney, R. M.; Olatoyinbo, P.; Yang, T.; Bassi, S.; Gonzales, V.; Chen, E.; Ashraf, F. B.; Roman, I. C.; Felix, J. R.; Alam, R. M.; Lay, J. A.; Johal, M. S.; Le Roch, K. G.; Tolstorukov, I.; Hernandez, J. B.; da Silva, F. L. B.; Lonardi, S.; Sazinsky, M. H.; Ray, A.

2025-10-31 bioinformatics
10.1101/2025.10.30.685465 bioRxiv
Show abstract

Scalable identification of antigen-specific antibodies from whole immune repertoire V(D)J sequences is a central challenge in biomedical engineering. We show that protein language models (PLMs) fine-tuned on antibody heavy-chain sequences can directly predict antigen specificity from unselected immune repertoires. We assessed our model, Antigen Specificity Predictor (ASPred), against SARS-CoV-2, influenza, and HIV-AIDS antigens, observing comparable predictive performance. In the whole immune repertoire V(D)J sequences of mice immunized with the SARS-CoV-2 spike proteins receptor-binding domain (RBD), ASPred identified antibody sequences specific to RBD. Several candidate sequences were validated, including one as a heavy chain-only nanobody with 20.7 nM dissociation constant. Molecular dynamics simulations supported the predicted interactions at coarse-grained and atomic levels. Benchmarking against Barcode-Enabled Antigen Mapping (BEAM) of B cell receptor sequence data had highly significant overlaps with ASPred predictions, suggesting scalability. The predicted SARS-CoV-2 binders differed substantially from training sequences, demonstrating generalization beyond sequence memorization. Together, we establish that heavy chain antibody sequences encode sufficient information for PLMs to infer specificity, offering a scalable framework for antibody discovery with broad applications.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.