New Protein Function Characterization for Human Paralog Discovery, Scraping the Bottom of the Genomics Barrel
BK, P.; Deng, W.; Jernigan, R. L.
Show abstract
Increasing the number of related protein paralogs is important for fully understanding protein relationships, yet it remains challenging for sequences in the twilight zone. Here, we present an integrated homolog detection framework that combines sequence-based (BLASTp, MMseqs2), structure-based (Foldseek), and embedding-distance-based (PROST) similarity metrics to identify additional paralogs. To characterize functionally related protein pairs, we develop protein-family-specific supervised logistic regression models trained on curated functional annotations from MEROPS proteases and KinHub kinases. The resulting model successfully classifies proteins, with ROC-AUC of 0.99 and F1-score of 0.92 for test datasets. Applying this model, we initially identify 686 protease and 298 kinase new candidates. Subsequent structural validation, and previous annotation comparisons yield 7 new protease and 3 new kinase paralogs in the human proteome, mostly lacking prior functional characterization. An additional outcome is structural identification of catalytically important residues for larger numbers of proteases and kinases. Despite the small number of new paralogs for well-studied proteases and kinases, our results demonstrate that integrating orthogonal homolog approaches with family-specific regression models provides a robust, scalable strategy for discovering new functionally related proteins, which is a generalizable approach for novel protein function discovery and can be applied more broadly to under-annotated proteomes.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 97%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 97%
- ECOD domain classification of 48 whole proteomes from AlphaFold Structure Database using DPAM 95%
Similar papers in this journal
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 97%
- Structural Coverage of the Human Interactome 96%
- A Unified Protein Embedding Model with Local and Global Structural Sensitivity 96%
Similar papers in this journal
- Unraveling cooperative and competitive interactions within protein triplets in the human interactome 96%
- Benchmarking Protein Language Models for Protein Crystallization 95%
- ProtAlign-ARG: Antibiotic Resistance Gene Characterization Integrating Protein Language Models and Alignment-Based Scoring 95%
Similar papers in this journal
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 96%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 95%
- Fast protein structure searching using structure graph embeddings 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.