ProTokens: Probabilistic Vocabulary for Compact and Informative Encodings of All-Atom Protein Structures
Lin, X.; Chen, Z.; Li, Y.; Ma, Z.; Fan, C.; Cao, Z.; Feng, S.; Gao, Y. Q.; Zhang, J.
Show abstract
Understanding functions of proteins and designing proteins committed to specific functions in silico are highly valuable for science, industry and therapeutics. However, there is a long-standing divergence in how to present function-related protein structures to the machine learning models: Although the 1-dimensional (1D) representation of proteins via Anfinsens tokens (i.e., amino acids) is sufficient in principle and more machine-friendly, it is less successful in structure-oriented protein design compared to symmetry-constrained 3D representation (i.e., atom coordinates). Aiming to bridge the gap between 1D and 3D protein representations and harvest the advantages of the two, we develop probabilistic tokenization theory for metastable protein structures. We present an unsupervised learning strategy, which conjugates inverse folding with structure prediction, to encode protein structures into artificial amino-acid tokens (ProTokens) and decode them back to atom coordinates. We show that tokenizing protein structures variationally can lead to compact and informative representations. Compared to amino acids -- the Anfinsens tokens -- ProTokens are easier to detokenize and more descriptive of finer conformational ensembles. Therefore, protein structures can be efficiently compressed, stored, aligned and compared in the form of ProTokens. By unifying the discrete and continuous representations of protein structures, ProTokens also enable all-atom protein structure design via various generative models without the concern of symmetry or modality mismatch, and allows scalable foundation models to perceive, process and explore the microscopic structures of biomolecules effectively.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- US-align: Universal Structure Alignments of Proteins, Nucleic Acids, and Macromolecular Complexes 96%
- Direct prediction of intrinsically disordered protein conformational properties from sequence 94%
- Predicting structures of large protein assemblies using combinatorial assembly algorithm and AlphaFold2 94%
Similar papers in this journal
- Improving the prediction of protein stability changes upon mutations by geometric learning and a pre-training strategy 96%
- A kinetic ensemble of the Alzheimer's Aβ peptide 94%
- Unconstrained generation of synthetic antibody-antigen structures to guide machine learning methodology for real-world antibody specificity prediction 94%
Similar papers in this journal
- CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon Aware Embeddings with Amino Acid Aware Embeddings and Wavelet-based Kolmogorov Arnold Network 94%
- Topology-Driven Negative Sampling Enhances Generalizability in Protein-Protein Interaction Prediction 94%
- Deep Local Analysis evaluates protein docking conformations with locally oriented cubes 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.