Robust Conditional Diffusion with Noisy Templates for Antibody Sequence-Structure Design
Liu, p.; Zhang, J.; Yan, C.
Show abstract
Antibodies specifically recognize antigens and play a central role in therapeutic discovery. Designing antibodies for a given antigen remains challenging because antigen-antibody complex data are limited, whereas the sequence and conformational spaces of complementarity-determining regions (CDRs) are large. Retrieved CDR templates from databases or candidate libraries can narrow the design space and improve controllability, but retrieval for novel antigens is often sparse and imperfect; treating retrieved templates as hard conditions can bias the denoising process and cause negative transfer. To address this problem, we propose Robust Conditional Diffusion with Noisy Templates for antibody sequence-structure design (NT-ABDiff), a joint diffusion framework that treats candidate CDR-only templates as optional and potentially unreliable conditions. NT-ABDiff uses reliability-aware template modulation to estimate the context-conditioned usefulness of each candidate and to adaptively reweight and fuse multiple templates during conditioning. We further train the model with mixed-quality and corrupted templates as conditional perturbation regularization, encouraging the denoiser to exploit informative templates while remaining stable when templates are uninformative. Experiments under controlled template shifts and a train-set retrieval evaluation show that NT-ABDiff improves CDR-H3 sequence recovery and structural accuracy over strong baselines, while retaining robustness to missing, mismatched, and corrupted templates. Under a stringent random-template CDR-H3 evaluation, NT-ABDiff improves amino-acid recovery (AAR) from 30.03% to 39.47% and reduces RMSD from 3.160 to 2.915A; with train-set retrieval candidates, it achieves 39.50% AAR and 2.76 {ring} A RMSD. Code, processed splits, {ring} configuration files, and evaluation scripts are available at https://github.com/ShiDeng7rz/NT-ABDiff.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 95%
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 94%
- TEINet: a deep learning framework for prediction of TCR-epitope binding specificity 94%
Similar papers in this journal
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 97%
- A curriculum learning approach to training antibody language models 96%
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.