Back

PepSeek: Universal Functional Peptide Discovery with Cooperation Between Specialized Deep Learning Models and Large Language Model

Gong, H.; Wang, Y.; Kong, Q.; Li, X.; Li, L.; Wan, B.; Zhao, Y.; Chen, G.; Chen, J.; Zhang, J.; Yu, Y.; Yang, X.; Zuo, X.; Li, Y.

2025-04-30 bioinformatics
10.1101/2025.04.29.641945 bioRxiv
Show abstract

Recent computational foundation models have revitalized the scientific discovery pipeline. However, developing foundational models for functional peptide discovery is costly due to the scarcity of wet-lab validated data. Meanwhile, conventional deep learning models are hard to generalize to unseen tasks or data distribution. Here, we introduce PepSeek, a universal approach for peptide discovery that synergistically integrates the most advanced large language model (LLM) with specialized small models. PepSeek harnesses the robust reasoning and generalization capabilities of LLM while leveraging the high predictive accuracy of specialized models trained for tasks such as antimicrobial activity regression and functional peptide generation. We have devised multiple collaborative strategies and task-specific modules demonstrating leading performance in peptide identification and generation. Notably, PepSeek achieves remarkable zero-shot prediction accuracy for peptides with diverse functionalities. We used PepSeek to identify a group of broad-spectrum antimicrobial peptide that exhibits low toxicity and high activity against drug-resistant bacteria, with the best surpassing all peptides currently undergoing clinical trials. Our framework establishes a new pipeline for scientific discovery with the help of LLM and specialized models.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.