pLM-Guided Inverse Folding for Antibody Sequence Design
Noske, V.; Koulischer, F.; Marchal, K.; Demeester, T.
Show abstract
Inverse folding, predicting amino acid sequences from three-dimensional structures, is a foundational task in computational protein design, yet it is hindered by the scarcity of structural data, which limits model training and risks overfitting. The standard approach fine-tunes general inverse folding models on domain-specific structural datasets like antibodies, but such data remain expensive. To enable inverse folders to benefit more from abundant sequence data, we propose combining ProteinMPNN, a general protein inverse folding model, with IgLM, an antibody-specific language model, via a training-free weighted ensemble of their predictions at inference time. Evaluated on antibody and nanobody structures, our results show that this approach substantially improves amino acid recovery over ProteinMPNN alone, approaching the performance of antibody-specific models like AntiFold while generating more diverse sequences. Even models already fine-tuned on antibody structures (AbMPNN) benefit from language model guidance, demonstrating that it complements structural fine-tuning and leads to more natural-looking sequences that still satisfy structural constraints.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Towards generalizable prediction of antibody thermostability using machine learning on sequence and structure features 94%
- AbDesign: Database of point mutants of antibodies with associated structures reveals poor generalization of binding predictions from machine learning models. 94%
- NAStructuralDB : Structural database to facilitate computational studies of molecular modeling and recognition of proteins with special focus on antibody-antigen interactions. 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.