Odyssey: reconstructing evolution through emergent consensus in the global proteome
Singhal, A.; Venkatasubramanian, S.; Moushegian, S.; Strutt, S.; Lin, M.; Lee, C.
Show abstract
We present Odyssey, a family of multimodal protein language models for sequence and structure generation, protein editing and design. We scale Odyssey to more than 102 billion parameters, trained over 1.1 x 1023 FLOPs. The Odyssey architecture uses context modalities, categorized as structural cues, semantic descriptions, and orthologous group metadata, and comprises two main components: a finite scalar quantizer for tokenizing continuous atomic coordinates, and a transformer stack for multimodal representation learning. Odyssey is trained via discrete diffusion, and characterizes the generative process as a time-dependent unmasking procedure. The finite scalar quantizer and transformer stack leverage the consensus mechanism, a replacement for attention that uses an iterative propagation scheme informed by local agreements between residues. Across various benchmarks, Odyssey achieves landmark performance for protein generation and protein structure discretization. Our empirical findings are supported by theoretical analysis.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Mapping the gene space at single-cell resolution with gene signal pattern analysis 95%
- Adversarial domain translation networks for fast and accurate integration of large-scale atlas-level single-cell datasets 94%
- Automated customization of large-scale spiking network models to neuronal population activity 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.