Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language Models
Li, F.-Z.; Amini, A. P.; Yue, Y.; Yang, K. K.; Lu, A. X.
Show abstract
Large pretrained protein language models (PLMs) have improved protein property and structure prediction from sequences via transfer learning, in which weights and representations from PLMs are repurposed for downstream tasks. Although PLMs have shown great promise, currently there is little understanding of how the features learned by pretraining relate to and are useful for downstream tasks. We perform a systematic analysis of transfer learning using PLMs, conducting 370 experiments across a comprehensive suite of factors including different downstream tasks, architectures, model sizes, model depths, and pretraining time. We observe that while almost all down-stream tasks do benefit from pretrained models compared to naive sequence representations, for the majority of tasks performance does not scale with pretraining, and instead relies on low-level features learned early in pretraining. Our results point to a mismatch between current PLM pretraining paradigms and most applications of these models, indicating a need for better pretraining methods.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences 98%
- Few-Shot Viral Variant Detection via Bayesian Active Learning and Biophysics 95%
- Predicting the unseen: a diffusion-based debiasing framework for transcriptional response prediction at single-cell resolution 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.