Scaling unlocks broader generation and deeper functional understanding of proteins
Bhatnagar, A.; Jain, S.; Beazer, J.; Curran, S. C.; Hoffnagle, A. M.; Ching, K.; Martyn, M.; Nayfach, S.; Ruffolo, J. A.; Madani, A.
Show abstract
Generative protein language models (PLMs) are powerful tools for designing proteins purpose-built to solve problems in medicine, agriculture, and industrial processes. Recent work has trained ever larger language models, but there has been little systematic study of the optimal training distributions and the influence of model scale on the sequences generated by PLMs. We introduce the ProGen3 family of sparse generative PLMs, and we develop compute-optimal scaling laws to scale up to a 46B-parameter model pre-trained on 1.5T amino acid tokens. Pro-Gen3s pre-training data is sampled from an optimized data distribution over the Profluent Protein Atlas v1, a carefully curated dataset of 3.4B full-length proteins. We evaluate for the first time in the wet lab the influence of model scale on the sequences generated by PLMs, and we find that larger models generate viable proteins for a much wider diversity of protein families. Finally, we find both computationally and experimentally that larger models are more responsive to alignment with laboratory data, resulting in improved protein fitness prediction and sequence generation capabilities. These results indicate that larger PLMs like ProGen3-46B trained on larger, well-curated datasets are powerful foundation models that push the frontier of protein design.1
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Sequence-based prediction of protein-protein interactions: a structure-aware interpretable deep learning model 96%
- An adversarial scheme for integrating multi-modal data on protein function 95%
- DynamicGT: a dynamic-aware geometric transformer model to predict protein binding interfaces in flexible and disordered regions 95%
Similar papers in this journal
- Adversarial domain translation networks for fast and accurate integration of large-scale atlas-level single-cell datasets 95%
- Mapping the gene space at single-cell resolution with gene signal pattern analysis 94%
- Biophysically Interpretable Inference of Cell Types from Multimodal Sequencing Data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.