Multi-Peptide Prompting Enables In-Context Learning in Protein Language
Almonte, J.; Vu, M. T.; Ahn, A.; Thiede, E. H.
Show abstract
Protein language models (PLMs) are trained primarily on individual protein sequences, yet many peptide-discovery problems require inference from only a small number of labeled examples. Here, we show that single-sequence PLMs can perform in-context peptide learning without gradient updates, task-specific retraining, or architectural modification. We introduce multi-peptide example prompts (MPEPs), in which demonstration peptides are concatenated with glycine spacers and used as context for scoring query peptides by their prompted probability. We evaluate this approach across a synthetic pattern-completion task, secondary-structure classification, and MHC-II binder prediction using both encoder-only ESM-2 models and decoder-only ProGen2 models. Across tasks, performance improves with the number of peptide examples and with model scale, indicating that PLMs can extract shared sequence-level properties from prompted examples. We further introduce a difference score that contrasts positive-example and negative-example MPEPs, reducing compositional biases in raw PLM probabilities and substantially improving classification. On MHC-II binder prediction, MPEP-based classification with larger ESM-2 models matches or exceeds low-data classifiers trained on frozen ESM-2 embeddings, while requiring no training. These results reveal an unexpected in-context inference capability in single-sequence PLMs and establish MPEP conditioning as a lightweight strategy for low-data peptide classification.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.