Accurately identifying nucleic-acid-binding sites through geometric graph learning on language model predicted structures
Song, Y.; Yuan, Q.; Zhao, H.; Yang, Y.
Show abstract
The interactions between nucleic acids and proteins are important in diverse biological processes. The high-quality prediction of nucleic-acid-binding sites continues to pose a significant challenge. Presently, the predictive efficacy of sequence-based methods is constrained by their exclusive consideration of sequence context information, whereas structure-based methods are unsuitable for proteins lacKing Known tertiary structures. Though protein structures predicted by AlphaFold2 could be used, the extensive computing requirement of AlphaFold2 hinders its use for genome-wide applications. Based on the recent breaKthrough of ESMFold for fast prediction of protein structures, we have developed GLMSite, which accurately identifies DNA and RNA-binding sites using geometric graph learning on ESMFold predicted structures. Here, the predicted protein structures are employed to construct protein structural graph with residues as nodes and spatially neighboring residue pairs for edges. The node representations are further enhanced through the pre-trained language model ProtTrans. The networK was trained using a geometric vector perceptron, and the geometric embeddings were subsequently fed into a common networK to acquire common binding characteristics. Then two fully connected layers were employed to learn specific binding patterns for DNA and RNA, respectively. Through comprehensive tests on DNA/RNA benchmarK datasets, GLMSite was shown to surpass the latest sequence-based methods and be comparable with structure-based methods. Moreover, the prediction was shown useful for the inference of nucleic-acid-binding proteins, demonstrating its potential for protein function discovery. The datasets, codes, together with trained models are available at https://github.com/biomed-AI/nucleic-acid-binding.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GraphGPSM: a global scoring model for protein structure using graph neural networks 98%
- A Reproducibility Analysis-based Statistical Framework for Residue-Residue Evolutionary Coupling Detection 98%
- Improved model quality assessment using sequence and structural information by enhanced deep neural networks 97%
Similar papers in this journal
- DeepUMQA: Ultrafast Shape Recognition-based Protein Model Quality Assessment using Deep Learning 97%
- Pair-EGRET: enhancing the prediction of protein-proteininteraction sites through graph attention networks and protein language models 97%
- FAPM: Functional Annotation of Proteins using Multi-Modal Models Beyond Structural Modeling 97%
Similar papers in this journal
- Pathfinder: protein folding pathway prediction based on conformational sampling 96%
- Hybridized distance- and contact-based hierarchical structure modeling for folding soluble and membrane proteins 96%
- Base-resolution prediction of transcription factor binding signals by a deep learning framework 95%
Similar papers in this journal
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 97%
- LinAliFold and CentroidLinAliFold: Fast RNA consensus secondary structure prediction for aligned sequences using beam search methods 96%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 96%
Similar papers in this journal
- Physical-aware model accuracy estimation for protein complex using deep learning method 98%
- SpatialPPI: three-dimensional space protein-protein interaction prediction with AlphaFold Multimer 96%
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.