StarFunc: fusing template-based and deep learning approaches for accurate protein function prediction
ZHANG, C.; Liu, Q.; Freddolino, L.
Show abstract
Deep learning has significantly advanced the development of high-performance methods for protein function prediction. Nonetheless, even for state-of-the-art deep learning approaches, template information remains an indispensable component in most cases. While many function prediction methods use templates identified through sequence homology or protein-protein interactions, very few methods detect templates through structural similarity, even though protein structures are the basis of their functions. Here, we describe our development of StarFunc, a composite approach that integrates state-of-the-art deep learning models seamlessly with template information from sequence homology, protein-protein interaction partners, proteins with similar structures, and protein domain families. Large-scale benchmarking and blind testing in the 5th Critical Assessment of Function Annotation (CAFA5) consistently demonstrate StarFuncs advantage when compared to both state-of-the-art deep learning methods and conventional template-based predictors.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CATHe: Detection of remote homologues for CATH superfamilies using embeddings from protein language models 96%
- CONSTRUCT: an algorithmic tool for identifying functional or structurally important regions in protein tertiary structure 96%
- SOLeNNoID: A Deep Learning Pipeline For Solenoid Residue Detection in Protein Structures 95%
Similar papers in this journal
- DeepSS2GO: protein function prediction from secondary structure 97%
- INTREPPPID - An Orthologue-Informed Quintuplet Network for Cross-Species Prediction of Protein-Protein Interaction 96%
- GraphCPLMQA: Assessing protein model quality based on deep graph coupled networks using protein language model 95%
Similar papers in this journal
Similar papers in this journal
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 95%
- Sensitive and error-tolerant annotation of protein-coding DNA with BATH 94%
- WAS IT A MATch I SAW? Approximate palindromes lead to overstated false match rates in benchmarks using reversed sequences 94%
Similar papers in this journal
- ChemPert: mapping between chemical perturbation and transcriptional response for non-cancer cells 95%
- The Protein Common Assembly Database (ProtCAD): A comprehensive structural resource of protein complexes 95%
- iModulonDB: a knowledgebase of microbial transcriptional regulation derived from machine learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.