Disulphide and sequence-encoded conformational priors guide nanobody structure prediction
Ali, M.; Jaskolowski, M.; Greenig, M.; Nguyen, T. H.; Smorodina, E.; Crnogaj, M.; Ramon, A. E.; Zhao, H.; Fernandez-Quintero, M. L.; Ghedin, E.; Greiff, V.; Sormanni, P.
Show abstract
Nanobody binding is largely governed by the HCDR3 loop, which adopts distinct placement regimes relative to the framework: compact, framework-contacting (kinked blueprint) and solvent-exposed (extended blueprint). Many nanobodies also contain additional cysteines that form non-canonical disulphide bonds, imposing covalent constraints on binding-loop conformations. Current structure predictors are typically trained and benchmarked with smooth coordinate-based objectives, so models may appear reasonable under root-mean-square deviation (RMSD), while adopting an incorrect HCDR3 blueprint or failing to recover the native disulphide connectivity, impacting paratope geometry and functional interpretation. Here, we show that the HCDR3 blueprint is predictable from sequence alone, allowing for explicit constraints during modelling. We implement these principles into NbForge, a lightweight nanobody folding model that incorporates blueprint- and disulphide-aware inductive biases and is trained with filtered self-distillation. NbForge improves recovery of HCDR3 blueprint and non-canonical disulphide formation over previous lightweight models and achieves coordinate accuracy at par to state-of-the-art, large, resource-intensive predictors, while running at sub-second inference speed. We show that using NbForge monomer models as templates further improves the success rate of predicting nanobody-antigen complexes. Together, these results motivate blueprint- and disulphide-aware benchmarks for nanobody modelling beyond RMSD, and show that appropriate inductive biases can close the performance gap to heavyweight predictors. We make the sequence classifier (NbFrame) and NbForge available for download and via a user-friendly web server.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Hierarchical design of multi-scale protein complexes by combinatorial assembly of oligomeric helical bundle and repeat protein building blocks 96%
- Improved protein structure refinement guided by deep learning based accuracy estimation 96%
- Deep-Learning Structure Elucidation from Single-Mutant Deep Mutational Scanning 95%
Similar papers in this journal
Similar papers in this journal
- Ig-VAE: Generative Modeling of Immunoglobulin Proteins by Direct 3D Coordinate Generation 95%
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 94%
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.