ANARCII: A Generalised Language Model for Antigen Receptor Numbering
Greenshields-Watson, A.; Agarwal, P.; Robinson, S. A.; Williams, B. H.; Gordon, G. L.; Capel, H. L.; Li, Y.; Spoendlin, F. C.; Boyles, F.; Deane, C. M.
Show abstract
Antigen receptor numbering allows the rapid delineation of the antigen-binding regions of antibody and T cell receptor (TCR) sequences, from sequence alone. It also allows the comparison of the vast diversity of antigen receptors in a consistent frame of reference. Numbering of antigen receptors is currently achieved by aligning sequences to a reference set. This approach may result in different numbering, depending on the reference set used or may fail to number query sequences derived from new species or rare sequence types. To address this problem, we have built a new numbering method (ANARCII) which requires no alignment step and is based on a Seq2Seq language model. Our results show that ANARCII can deal with the complexity that arises in experimentally collected sequencing data and generalise to sequences which are highly dissimilar to those in training. In test sets designed to contain challenging and ambiguous sequence patterns ANARCII numbering was identical to existing methods for over 99.99% of conserved residues and over 99.94% for complete CDR regions. The lightweight architecture allows numbering of over 90,000 sequences per minute on a single A100 GPU. Furthermore, the ANARCII package can be conditioned to fit rare sequence types and provide new training data for fine-tuning. We demonstrate that fine-tuned versions of ANARCII can correctly number other immunoglobulin domains such as TCRs and VNARs. Our model is freely available as a web tool (https://github.com/oxpig/ANARCII), as well as a package for high throughput numbering of next generation sequencing data (https://opig.stats.ox.ac.uk/webapps/sabdab-sabpred/sabpred/anarcii/).
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ParaSurf: A Surface-Based Deep Learning Approach for Paratope-Antigen Interaction Prediction 95%
- STCRpy: a software suite for T cell receptor structure parsing, interaction profiling and machine learning dataset preparation 95%
- ProtMamba: a homology-aware but alignment-free protein state space model 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Humatch - fast, gene-specific joint humanisation of antibody heavy and light chains 94%
- BioPhi: A platform for antibody design, humanization and humanness evaluation based on natural antibody repertoires and deep learning 93%
- AbDesign: Database of point mutants of antibodies with associated structures reveals poor generalization of binding predictions from machine learning models. 93%
Similar papers in this journal
- Interpretable deep learning to uncover the molecular binding patterns determining TCR-epitope interactions 94%
- Guiding a language-model based protein design method towards MHC Class-I immune-visibility profiles for vaccines and therapeutics 94%
- DoggifAI: a transformer based approach for antibodycaninisation 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.