PepBERT: Lightweight language models for peptide representation
Du, Z.; Li, Y.
Show abstract
Protein language models (pLMs) have been widely adopted for various protein and peptide-related downstream tasks and demonstrated promising performance. However, short peptides are significantly underrepresented in commonly used pLM training datasets. For example, only 2.8% of sequences in the UniProt Reference Cluster (UniRef) contain fewer than 50 residues, which potentially limits the effectiveness of pLMs for peptide-specific applications. Here, we present PepBERT, a lightweight and efficient peptide language model specifically designed for encoding peptide sequences. Two versions of the model--PepBERT-large (4.9 million parameters) and PepBERT-small (1.86 million parameters)--were pretrained from scratch using four custom peptide datasets and evaluated on nine peptide-related downstream prediction tasks. Both PepBERT models achieved performance superior to or comparable to the benchmark model, ESM-2 with 7.5 million parameters, on 8 out of 9 datasets. Overall, PepBERT provides a compact yet effective solution for generating high-quality peptide representations for downstream applications. By enabling more accurate representation and prediction of bioactive peptides, PepBERT can accelerate the discovery of food-derived bioactive peptides with health-promoting properties, supporting the development of sustainable functional foods and value-added utilization of food processing by-products. The datasets, source codes, pretrained models, and tutorials for the usage of PepBERT are available at https://github.com/dzjxzyd/PepBERT.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Peptriever: A Bi-Encoder approach for large-scale protein-peptide binding search 96%
- DistilProtBert: A distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts 95%
- BERTMHC: Improves MHC-peptide class II interaction prediction with transformer and multiple instance learning 95%
Similar papers in this journal
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 94%
- AI-Guided Discovery and Optimization of Antimicrobial Peptides Through Species-Aware Language Model 94%
- ProDualNet: Dual-Target Protein Sequence Design Method Based on Protein Language Model and Structure Model 94%
Similar papers in this journal
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 93%
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 93%
- ProteinPrompt: a webserver for predicting protein-protein interactions 92%
Similar papers in this journal
- Accelerating protein engineering with fitness landscape modeling and reinforcement learning 92%
- Improving protein function prediction with synthetic feature samples created by generative adversarial networks 91%
- Evaluating generalizability of artificial intelligence models for molecular datasets 91%
Similar papers in this journal
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 94%
- G-PLIP: Knowledge graph neural network for structure-free protein-ligand bioactivity prediction 92%
- Towards mechanistic models of mutational effects: Deep Learning on Alzheimer's Aβ peptide 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.