PeptideMTR: Scaling SMILES-Based Language Models for Therapeutic Peptide Engineering
Feller, A. L.; Secor, M.; Swanson, S.; Wilke, C. O.; Deibler, K.
Show abstract
Therapeutic peptides occupy a unique middle ground in drug discovery, offering the high specificity of protein interactions with the chemical diversity of small molecules, yet they currently fall in a computational blind spot. Existing AI tools cannot handle them effectively: protein models are restricted to natural amino acids, while chemical models struggle to process large, polymer-like sequences. This disconnect has forced the field to rely on static chemical descriptors that fail to capture subtle chemical details. To bridge this gap, we present PeptideCLM-2, a chemical language model trained on over 100 million molecules to natively represent complex peptide chemistry. PeptideCLM-2 consistently outperforms both chemical descriptors and specialized AI models on critical drug development tasks, including aggregation, membrane diffusion, and cell targeting. Notably, we find that when model parameters reach the 100 million scale, the transformer architecture is able to learn chemical properties from molecular syntax alone.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PepFlow: direct conformational sampling from peptide energy landscapes through hypernetwork-conditioned diffusion 96%
- Annotating metabolite mass spectra with domain-inspired chemical formula transformers 96%
- Accelerating protein engineering with fitness landscape modeling and reinforcement learning 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.