Beware of Data Leakage from Protein LLM Pretraining
Hermann, L.; Fiedler, T.; Nguyen, H. A.; Nowicka, M.; Bartoszewicz, J. M.
Show abstract
Pretrained protein language models are becoming increasingly popular as a backbone for protein property inference tasks such as structure prediction or function annotation, accelerating biological research. However, related research oftentimes does not consider the effects of data leakage from pretraining on the actual downstream task, resulting in potentially unrealistic performance estimates. Reported generalization might not necessarily be reproducible for proteins highly dissimilar from the pretraining set. In this work, we measure the effects of data leakage from protein language model pretraining in the domain of protein thermostability prediction. Specifically, we compare two different dataset split strategies: a pretraining-aware split, designed to avoid similarity between pretraining data and the held-out test sets, and a commonly-used naive split, relying on clustering the training data for a downstream task without taking the pretraining data into account. Our experiments suggest that data leakage from language model pretraining shows consistent effects on melting point prediction across all experiments, distorting the measured performance. The source code and our dataset splits are available at https://github.com/tfiedlerdev/pretraining-aware-hotprot.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FAPM: Functional Annotation of Proteins using Multi-Modal Models Beyond Structural Modeling 96%
- Pair-EGRET: enhancing the prediction of protein-proteininteraction sites through graph attention networks and protein language models 96%
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 96%
Similar papers in this journal
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 96%
- Prop3D: A Flexible, Python-based Platform for Machine Learning with Protein Structural Properties and Biophysical Data 95%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 95%
Similar papers in this journal
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 97%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 96%
- KSMoFinder - Knowledge graph embedding of proteins and motifs for predicting kinases of human phosphosites 95%
Similar papers in this journal
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 96%
- Designing diverse and high-performance proteins with a large language model in the loop 96%
- Protein Stability Prediction by Fine-tuning a Protein Language Model on a Mega-scale Dataset 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.