Machine Learning-Guided Antibody Engineering That Leverages Domain Knowledge To Overcome The Small Data Problem
Clark, T.; Subramanian, V.; Jayaraman, A.; Fitzpatrick, E.; Gopal, R.; Pentakota, N.; Rurak, T.; Anand, S.; Viglione, A.; Tharakaraman, K.; Raman, R.; Sasisekharan, R.
Show abstract
The application of Machine Learning (ML) tools to engineer novel antibodies having predictable functional properties is gaining prominence. Herein, we present a platform that employs an ML-guided optimization of the complementarity-determining region (CDR) together with a CDR framework (FR) shuffling method to engineer affinity-enhanced and clinically developable monoclonal antibodies (mAbs) from a limited experimental screen space (order of 10^2 designs) using only two experimental iterations. Although high-complexity deep learning models like graph neural networks (GNNs) and large language models (LLMs) have shown success on protein folding with large dataset sizes, the small and biased nature of the publicly available antibody-antigen interaction datasets is not sufficient to capture the diversity of mutations virtually screened using these models in an affinity enhancement campaign. To address this key gap, we introduced inductive biases learned from extensive domain knowledge on protein-protein interactions through feature engineering and selected model hyper parameters to reduce overfitting of the limited interaction datasets. Notably we show that this platform performs better than GNNs and LLMs on an in-house validation dataset that is enriched in diverse CDR mutations that go beyond alanine-scanning. To illustrate the broad applicability of this platform, we successfully solved a challenging problem of redesigning two different anti-SARS-COV-2 mAbs to enhance affinity (up to 2 orders of magnitude) and neutralizing potency against the dynamically evolving SARS-COV-2 Omicron variants.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AlphaBind, a Domain-Specific Model to Predict and Optimize Antibody-Antigen Binding Affinity 98%
- Tuning antibody stability and function by rational designs of framework mutations 96%
- Towards generalizable prediction of antibody thermostability using machine learning on sequence and structure features 96%
Similar papers in this journal
- Machine Learning Optimization of Candidate Antibodies Yields Highly Diverse Sub-nanomolar Affinity Antibody Libraries 96%
- Selection, biophysical and structural analysis of synthetic nanobodies that effectively neutralize SARS-CoV-2 95%
- Sequence signatures of two IGHV3-53/3-66 public clonotypes to SARS-CoV-2 receptor binding domain 95%
Similar papers in this journal
- Peptide-antibody Fusions Engineered by Phage Display Exhibit Ultrapotent and Broad Neutralization of SARS-CoV-2 Variants 96%
- Rationalizing diverse binding mechanisms to the same protein fold: in-sights for ligand recognition and biosensor design 94%
- Integrative x-ray structure and molecular modeling for the rationalization of procaspase-8 inhibitor potency and selectivity 94%
Similar papers in this journal
- Enhanced Sequence-Activity Mapping and Evolution of Artificial Metalloenzymes by Active Learning 94%
- A High-Throughput Screen Reveals the Structure-Activity Relationship of the Antimicrobial Lasso Peptide Ubonodin 94%
- In vitro selection of cyclized, glycosylated peptide antigens that tightly bind HIV high mannose patch antibodies 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.