Separating selection from mutation in antibody language models
Matsen, F. A.; Dumm, W.; Sung, K.; Johnson, M. M.; Rich, D.; Starr, T.; Song, Y. S.; Fukuyama, J.; Haddox, H. K.
Show abstract
Antibodies are encoded by nucleotide sequences that are generated by V(D)J recombination and evolve according to mutation and selection processes. Existing antibody language models, however, focus exclusively on antibodies as strings of amino acids and are fitted using standard language modeling objectives such as masked or autoregressive prediction. In this paper, we first show that fitting models using this objective implicitly incorporates nucleotide-level mutation processes as part of the protein language model, which degrades performance when predicting effects of mutations on functional properties of antibodies. To address this limitation, we devise a new framework: a Deep Amino acid Selection Model (DASM) that learns the selection effects of amino-acid mutations while explicitly factoring out the nucleotide-level mutation process. By fitting selection as a separate term from the mutation process, the DASM exclusively quantifies functional effects: effects that change some aspect of the function of the antibody. This factorization leads to substantially improved performance on standard functional benchmarks. Moreover, our model is an order of magnitude smaller and multiple orders of magnitude faster to evaluate than existing approaches, as well as being readily interpretable.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Nucleotide context models outperform protein language models for predicting antibody affinity maturation 98%
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 98%
- Population variability in the generation and thymic selection of T-cell repertoires 96%
Similar papers in this journal
Similar papers in this journal
- Designing meaningful continuous representations of T cell receptor sequences with deep generative models 97%
- Protein Sequence Modelling with Bayesian Flow Networks 96%
- Hi-C-LSTM: Learning representations of chromatin contacts using a recurrent neural network identifies genomic drivers of conformation 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.