AbAgym: a well-curated dataset for the mutational analysis of antibody-antigen complexes
Cia, G.; Li, D.; Poblete, S.; Rooman, M.; Pucci, F.
Show abstract
With monoclonal antibodies becoming one of the largest classes of biopharmaceuticals, it is important to have curated data to train computational models that can accelerate their design. Despite the massive amount of mutagenesis data generated on antibody-antigen interactions, only a few small, well-curated datasets are available. This paper introduces AbAgym, a manually curated dataset comprising approximately 335k mutations in antibody-antigen complexes, including one tenth of interface mutations, whose effects on antibody-antigen binding have been experimentally quantified through deep mutational scanning (DMS) experiments. We collected and curated 67 DMS datasets from the literature together with the three-dimensional structure of each antibody-antigen complex. We benchmarked the performance of established force field methods as well as recent machine learning models that predict the change in binding affinity upon mutation. The former achieved modest performance, whereas the latter performed only marginally better than random. Finally, our analysis of hotspot residues responsible for immune evasion highlights the importance of accounting for biological complexities, such as conformational changes or oligomeric states that influence antibody-antigen binding, which are often overlooked. Abagym is freely available for academic use at https://github.com/3BioCompBio/Abagym.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Exploring the Potential of Structure-Based Deep Learning Approaches for T cell Receptor Design 96%
- From complete cross-docking to partners identification and binding sites predictions 96%
- Computer-guided Binding Mode Identification and Affinity Improvement of an LRR Protein Binder without Structure Determination 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.