ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention
Li, M.; Tan, Y.; Ma, X.; Zhong, B.; Zhou, Z.; Yu, H.; Ouyang, W.; Hong, L.; Zhou, B.; Tan, P.
Show abstract
Protein language models (PLMs) have shown remarkable capabilities in various protein function prediction tasks. However, while protein function is intricately tied to structure, most existing PLMs do not incorporate protein structure information. To address this issue, we introduce ProSST, a Transformer-based protein language model that seamlessly integrates both protein sequences and structures. ProSST incorporates a structure quantization module and a Transformer architecture with disentangled attention. The structure quantization module translates a 3D protein structure into a sequence of discrete tokens by first serializing the protein structure into residue-level local structures and then embeds them into dense vector space. These vectors are then quantized into discrete structure tokens by a pre-trained clustering model. These tokens serve as an effective protein structure representation. Furthermore, ProSST explicitly learns the relationship between protein residue token sequences and structure token sequences through the sequence-structure disentangled attention. We pre-train ProSST on millions of protein structures using a masked language model objective, enabling it to learn comprehensive contextual representations of proteins. To evaluate the proposed ProSST, we conduct extensive experiments on the zero-shot mutation effect prediction and several supervised downstream tasks, where ProSST achieves the state-of-the-art performance among all baselines. Our code and pretrained models are publicly available 2.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 98%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 98%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 96%
Similar papers in this journal
- ProALIGN: Directly learning alignments for protein structure prediction via exploiting context-specific alignment motifs 97%
- Combined topological data analysis and geometric deep learning reveal niches by the quantification of protein binding pockets 96%
- Critiquing Protein Family Classification Models Using Sufficient Input Subsets 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.