Back

Language Modeling Materializes a World Model of Protein Biology

Candido, S.; Hayes, T.; Derry, A.; Rao, R.; Lin, Z.; Verkuil, R.; Wu, B. Z.; Lee, J. S.; Bruguera, E. S.; Keval, J. A.; Kopylov, M.; Pak, J. E.; Wu, W.; Thomas, N.; Mataraso, S.; Hsu, A.; Trotman-Grant, A. C.; Fatras, K.; dos Santos Costa, A.; Badkundri, R.; Akin, H.; Oktay, D.; Deaton, J.; Montabana, E.; Sitwala, H.; Yu, Y.; Wiggert, M.; Carlin, D. A.; Goering, A. W.; Blazejewski, T.; Sandora, M.; Hla, M.; Jia, T. Z.; Kloker, L. H.; Sofroniew, N. J.; Uehara, M.; Pannu, J.; Bachas, S.; Liu, D. S.; Sercu, T.; Rives, A.

2026-06-04 bioinformatics
10.64898/2026.06.03.729735 bioRxiv
Show abstract

Proteins are fundamental to life. The full extent of their biology is beyond our ability to characterize with experimental approaches in the physical laboratory. Accurate digital representations could accelerate the discovery of protein biology through virtual experiments. We propose language modeling to learn unified and general representations that can be scaled to all of protein biology. Building on these representations, we develop a structure prediction model that exceeds the performance of established methods for biomolecular complex prediction across benchmarks, including for the interactions of antibodies with their targets. A simple search procedure yields high experimental success rates for the discovery of proteins with nanomolar binding affinities for both miniproteins and single-chain antibodies, a modality critical for therapeutic design. Study of the concepts in the language models representation space reveals a systematic organization aligned with the reductionist understanding of proteins developed through empirical science. Leveraging this organization, we generate a comprehensive map of protein biology encompassing over 6.8 billion sequences and 1.1 billion predicted structures, identifying connections across known and unknown biology. As a whole, this shows language modeling as a powerful substrate for representing the biology of proteins, operating across scales from the prediction and design of protein interactions at the atomic level, to identifying properties of proteins at different levels of granularity and abstraction, to the scale of mapping connections between proteins across billions of years of evolution.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.