Language Modeling Materializes a World Model of Protein Biology
Candido, S.; Hayes, T.; Derry, A.; Rao, R.; Lin, Z.; Verkuil, R.; Wu, B. Z.; Lee, J. S.; Bruguera, E. S.; Keval, J. A.; Kopylov, M.; Pak, J. E.; Wu, W.; Thomas, N.; Mataraso, S.; Hsu, A.; Trotman-Grant, A. C.; Fatras, K.; dos Santos Costa, A.; Badkundri, R.; Akin, H.; Oktay, D.; Deaton, J.; Montabana, E.; Sitwala, H.; Yu, Y.; Wiggert, M.; Carlin, D. A.; Goering, A. W.; Blazejewski, T.; Sandora, M.; Hla, M.; Jia, T. Z.; Kloker, L. H.; Sofroniew, N. J.; Uehara, M.; Pannu, J.; Bachas, S.; Liu, D. S.; Sercu, T.; Rives, A.
Show abstract
Proteins are fundamental to life. The full extent of their biology is beyond our ability to characterize with experimental approaches in the physical laboratory. Accurate digital representations could accelerate the discovery of protein biology through virtual experiments. We propose language modeling to learn unified and general representations that can be scaled to all of protein biology. Building on these representations, we develop a structure prediction model that exceeds the performance of established methods for biomolecular complex prediction across benchmarks, including for the interactions of antibodies with their targets. A simple search procedure yields high experimental success rates for the discovery of proteins with nanomolar binding affinities for both miniproteins and single-chain antibodies, a modality critical for therapeutic design. Study of the concepts in the language models representation space reveals a systematic organization aligned with the reductionist understanding of proteins developed through empirical science. Leveraging this organization, we generate a comprehensive map of protein biology encompassing over 6.8 billion sequences and 1.1 billion predicted structures, identifying connections across known and unknown biology. As a whole, this shows language modeling as a powerful substrate for representing the biology of proteins, operating across scales from the prediction and design of protein interactions at the atomic level, to identifying properties of proteins at different levels of granularity and abstraction, to the scale of mapping connections between proteins across billions of years of evolution.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Multi-scale classification decodes the complexity of the human E3 ligome 97%
- Understanding epistatic networks in the B1 -lactamases through coevolutionary statistical modeling and deep mutational scanning 96%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.