Understanding Language Model Scaling on Protein Fitness Prediction
Hou, C.; Liu, D.; Zafar, A.; Shen, Y.
Show abstract
Protein language models, and models that incorporate structure or homologous sequences, estimate sequence likelihoods p(sequence) that reflect the protein fitness landscape and are commonly used in mutation effect prediction and protein design. It is widely believed in deep learning field that larger model performs better across tasks. However, for fitness prediction, language model performance declines beyond a certain size, raising concerns about their scalability. Here, we showed that model size, training dataset, and stochastic elements can bias the predicted p(sequence) away from real fitness. Model performance on fitness prediction depends on how well p(sequence) matches evolutionary patterns in homologs, which is best achieved at a moderate p(sequence) level for most proteins. At extreme predicted wild-type sequence likelihoods, models predict uniformly low or high likelihoods for nearly all mutations, failing to reflect the real fitness landscape. Notably, larger models tend to predict proteins with higher p(sequence), which may exceed the moderate range and thus reduce performance. Our findings clarify the scaling behavior of protein models on fitness prediction and provide practical guidelines for their application and future development.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Generalizable and scalable protein stability prediction with rewired protein generative models 96%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 96%
- Understanding epistatic networks in the B1 -lactamases through coevolutionary statistical modeling and deep mutational scanning 96%
Similar papers in this journal
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 96%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 96%
- The amino acid sequence determines protein abundance through its conformational stability and reduced synthesis cost. 95%
Similar papers in this journal
- Expert-guided protein Language Models enable accurate and blazingly fast fitness prediction 95%
- BindPred: A Framework for Predicting Protein-Protein Binding Affinity from Language Model Embeddings 94%
- Deep Local Analysis deconstructs protein-protein interfaces and accurately estimates binding affinity changes upon mutation 94%
Similar papers in this journal
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 94%
- Inferring protein fitness landscapes from laboratory evolution experiments 93%
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.