Attention is all you need for general-purpose protein structure embedding
Cui, X.
Show abstract
General-purpose protein structure embedding can be used for many important protein biology tasks, such as protein design, drug design and binding affinity prediction. Recent researches have shown that attention-based encoder layers are more suitable to learn high-level features. Based on this key observation, we treat low-level representation learning and high-level representation learning separately, and propose a two-level general-purpose protein structure embedding neural network, called ContactLib-ATT. On the local embedding level, a simple yet meaningful hydrogen-bond representation is learned. On the global embedding level, attention-based encoder layers are employed for global representation learning. In our experiments, ContactLib-ATT achieves a SCOP superfamily classification accuracy of 82.4% (i.e., 6.7% higher than state-of-the-art method) on the SCOP40 2.07 dataset. Moreover, ContactLib-ATT is demonstrated to successfully simulate a structure-based search engine for remote homologous proteins, and our top-10 candidate list contains at least one remote homolog with a probability of 91.9%. Source codes: https://github.com/xfcui/contactlib.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- A Gated Graph Transformer for Protein ComplexStructure Quality Assessment and its Performancein CASP15 98%
- DeepUMQA: Ultrafast Shape Recognition-based Protein Model Quality Assessment using Deep Learning 97%
- A deep dilated convolutional residual network for predicting interchain contacts of protein homodimers 97%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Structure-aware Protein Solubility Prediction From Sequence Through Graph Convolutional Network And Predicted Contact Map 96%
- All-Atom Protein Sequence Design using Discrete Diffusion Models 94%
- Chemical Genomics Language Model toward Reliable and Explainable Compound-Protein Interaction Exploration 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.