Fast and accurate protein function prediction from sequence through pretrained language model and homology-based label diffusion
Yuan, Q.; Xie, J.; Xie, J.; Zhao, H.; Yang, Y.
Show abstract
Protein function prediction is an essential task in bioinformatics which benefits disease mechanism elucidation and drug target discovery. Due to the explosive growth of proteins in sequence databases and the diversity of their functions, it remains challenging to fast and accurately predict protein functions from sequences alone. Although many methods have integrated protein structures, biological networks or literature information to improve performance, these extra features are often unavailable for most proteins. Here, we propose SPROF-GO, a Sequence-based alignment-free PROtein Function predictor which leverages a pretrained language model to efficiently extract informative sequence embeddings and employs self-attention pooling to focus on important residues. The prediction is further advanced by exploiting the homology information and accounting for the overlapping communities of proteins with related functions through the label diffusion algorithm. SPROF-GO was shown to surpass state-of-the-art sequence-based and even network-based approaches by more than 14.5%, 27.3% and 10.1% in AUPR on the three sub-ontology test sets, respectively. Our method was also demonstrated to generalize well on non-homologous proteins and unseen species. Finally, visualization based on the attention mechanism indicated that SPROF-GO is able to capture sequence domains useful for function prediction. Key pointsO_LISPROF-GO is a sequence-based protein function predictor which leverages a pretrained language model to efficiently extract informative sequence embeddings, thus bypassing expensive database searches. C_LIO_LISPROF-GO employs self-attention pooling to capture sequence domains useful for function prediction and provide interpretability. C_LIO_LISPROF-GO applies hierarchical learning strategy to produce consistent predictions and label diffusion to exploit the homology information. C_LIO_LISPROF-GO is accurate and robust, with better performance than state-of-the-art sequence-based and even network-based approaches, and great generalization ability on non-homologous proteins and unseen species C_LI
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 96%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 96%
- scTensor detects many-to-many cell-cell interactions from single cell RNA-sequencing data 95%
Similar papers in this journal
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 99%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 96%
- KSMoFinder - Knowledge graph embedding of proteins and motifs for predicting kinases of human phosphosites 95%
Similar papers in this journal
- A Generalized Higher-order Correlation Analysis Framework for Multi-Omics Network Inference 96%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 95%
- Explainable deep transfer learning model for disease risk prediction using high-dimensional genomic data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.