Improving nanobody structure prediction with self-distillation
Ali, M.; Greenig, M.; Jaskolowski, M.; Crnogaj, M.; Smorodina, E.; Zhao, H.; Greiff, V.; Sormanni, P.
Show abstract
Nanobodies are increasingly attractive therapeutic and biotechnological molecules, yet accurate structure prediction of their highly variable H-CDR3 loops remains a central challenge for machine learning models. Here, we investigate whether nanobody-specific structure prediction can be improved through curated synthetic data strategies. We systematically evaluate different data augmentation regimes, including self-distillation from unlabelled VHH sequences. To ensure structural plausibility of synthetic training samples, we develop NanoKink, the first sequence-based classifier of kinked versus extended H-CDR3 conformations, and apply stringent filtering criteria for non-canonical disulfide bond placement and confor-mational accuracy. On a curated benchmark enriched for challenging nanobody features, we show that, for a fixed training compute budget, a nanobody-specific model trained with filtered synthetic data significantly improves over baseline models and NanobodyBuilder2, achieving lower mean H-CDR3 RMSD and fewer structural violations, while remaining competitive with AlphaFold3 at approximately two orders of magnitude lower per-structure inference time. Our results highlight promising directions in synthetic data generation for nanobody structure modelling and provide a practical framework for optimisation of VHH structure prediction models.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Ab-Ligity: Identifying sequence-dissimilar antibodies that bind to the same epitope 95%
- AbDesign: Database of point mutants of antibodies with associated structures reveals poor generalization of binding predictions from machine learning models. 95%
- NAStructuralDB : Structural database to facilitate computational studies of molecular modeling and recognition of proteins with special focus on antibody-antigen interactions. 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.