Back

Improving nanobody structure prediction with self-distillation

Ali, M.; Greenig, M.; Jaskolowski, M.; Crnogaj, M.; Smorodina, E.; Zhao, H.; Greiff, V.; Sormanni, P.

2025-12-02 bioengineering
10.64898/2025.12.01.691162 bioRxiv
Show abstract

Nanobodies are increasingly attractive therapeutic and biotechnological molecules, yet accurate structure prediction of their highly variable H-CDR3 loops remains a central challenge for machine learning models. Here, we investigate whether nanobody-specific structure prediction can be improved through curated synthetic data strategies. We systematically evaluate different data augmentation regimes, including self-distillation from unlabelled VHH sequences. To ensure structural plausibility of synthetic training samples, we develop NanoKink, the first sequence-based classifier of kinked versus extended H-CDR3 conformations, and apply stringent filtering criteria for non-canonical disulfide bond placement and confor-mational accuracy. On a curated benchmark enriched for challenging nanobody features, we show that, for a fixed training compute budget, a nanobody-specific model trained with filtered synthetic data significantly improves over baseline models and NanobodyBuilder2, achieving lower mean H-CDR3 RMSD and fewer structural violations, while remaining competitive with AlphaFold3 at approximately two orders of magnitude lower per-structure inference time. Our results highlight promising directions in synthetic data generation for nanobody structure modelling and provide a practical framework for optimisation of VHH structure prediction models.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.