BOND-PEP: topology-conditioned bipartite alignment for evidence-grounded peptide binder generation
Ding, W.
Show abstract
Peptide binders can modulate proteins that remain challenging for small molecules, but discovering high-affinity, selective peptides is still slow and sample-intensive. Sequence-first generators could scale design when structures are unavailable or conformationally heterogeneous, yet they often trade diversity for control: unconstrained sampling is inefficient while conditioning remains largely implicit. This limitation is exacerbated by the uneven transfer of protein language model priors to short peptides. Here we present BOND-PEP, a retrieval-augmented, bipartite-aligned, topology-conditioned framework that converts empirical binding evidence into an explicit, residue-resolved conditioning state for peptide generation. BOND-PEP shows low perplexity together with satisfactory free-generation hit rates and sequence novelty under a fair evaluation protocol and decoding budget. Compared with existing peptide generation methods, BOND-PEP achieves state-of-the-art results that match or improve upon validated peptide-protein sequence pairs. In total, BOND-PEP provides a practical, sequence-only route to controllable de novo peptide binder generation under noisy labels and distribution shift.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Interpreting Neural Networks for Biological Sequences by Learning Stochastic Masks 95%
- PSICHIC: physicochemical graph neural network for learning protein-ligand interaction fingerprints from sequence data 95%
- PepFlow: direct conformational sampling from peptide energy landscapes through hypernetwork-conditioned diffusion 95%
Similar papers in this journal
Similar papers in this journal
- Undersampling and the inference of coevolution in proteins 95%
- DynamicGT: a dynamic-aware geometric transformer model to predict protein binding interfaces in flexible and disordered regions 95%
- Engineering of highly active and diverse nuclease enzymes by combining machine learning and ultra-high-throughput screening 94%
Similar papers in this journal
- Multi-omics integration and regulatory inference for unpaired single-cell data with a graph-linked unified embedding framework 94%
- Single-sequence protein structure prediction using language models from deep learning 94%
- Probing molecular specificity with deep sequencing and biophysically interpretable machine learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.