Protein Language Models: Is Scaling Necessary?
Fournier, Q.; Vernon, R. M.; van der Sloot, A.; Schulz, B.; Chandar, S.; Langmead, C. J.
Show abstract
Public protein sequence databases contain samples from the fitness landscape explored by nature. Protein language models (pLMs) pre-trained on these sequences aim to capture this landscape for tasks like property prediction and protein design. Following the same trend as in natural language processing, pLMs have continuously been scaled up. However, the premise that scale leads to better performance assumes that source databases provide an accurate representation of the underlying fitness landscape, which is likely false. By developing an efficient codebase, designing a modern architecture, and addressing data quality concerns such as sample bias, we introduce AMPLIFY, a best-in-class pLM that is orders of magnitude less expensive to train and deploy than previous models. Furthermore, to support the scientific community and democratize the training of pLMs, we have open-sourced AMPLIFYs pre-training codebase, data, and model checkpoints.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DistilProtBert: A distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts 97%
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 96%
- ProteinBERT: A universal deep-learning model of protein sequence and function 96%
Similar papers in this journal
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 96%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 95%
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.