Large Numbers of New Human Paralogs Discovered
Bk, P.; Deng, W.; Hemmat, M. A.; Daniels, I. L.; Jernigan, R. L.
Show abstract
The identification of paralogs is critical for understanding protein evolution, function, and for drug design, yet many human proteins remain unannotated and poorly classified. Sequence-based homology detection alone often fails to detect distant paralogs, especially in the "twilight zone" or beyond, regarding sequence identity. Here we present an integrated homolog detection framework that combines results from BLASTp, MMseqs2, Foldseek, and the large protein language model-based tool PROST, followed by validation based on comparison of structures and for enzymes comparison of the specific structures of the catalytic residues. Using all-versus-all exhaustive comparisons across the 20,647 human proteins, we systematically identify novel paralogs and assess their catalytic residues for two serine protease clans. We discovered 14 previously uncharacterized human serine carboxypeptidases, validated against experimentally determined PDB structures, with 11 of these displaying conserved catalytic triads. We further identify 203 new paralogs for human kinases, with 163 of these in the major clusters that represent previously uncharacterized kinase subtypes and 30 putative novel human transcription factors. Across both serine protease subtypes, structural alignments enable the prediction of the previously unknown catalytic residues for those lacking UniProt annotations of active site residues. By integrating sequence, structure, and LPLM embedding-based approaches, the framework enables the discovery of surprisingly large numbers of unknown paralogs, permitting defining catalytic residues, and expands the understanding of protein functional landscapes. These findings provide the foundation for a large number of future functional, evolutionary, and therapeutic investigations.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Fast Approach for Structural and Evolutionary Analysis Based on Energetic Profile Protein Comparison 96%
- Mapping specificity, entropy, allosteric changes and substrates in blood proteases by a high-throughput protease screen 95%
- High-accuracy protein complex structure modeling based on sequence-derived structure complementarity 94%
Similar papers in this journal
- Distribution of disease-causing germline mutations in coiled-coils suggests essential role of their N-terminal region 93%
- Evaluating the Significance of Embedding-Based Protein Sequence Alignment with Clustering and Double Dynamic Programming for Remote Homology 93%
- Two sequence- and two structure-based ML models have learned different aspects of protein biochemistry 93%
Similar papers in this journal
- Prediction of protein assemblies by structure sampling followed by interface-focused scoring 94%
- Novel sampling strategies and a coarse-grained score function for docking homomers, flexible heteromers, and oligosaccharides using Rosetta in CAPRI Rounds 37-45 94%
- Do Newly Born Orphan Proteins Resemble Never Born Proteins? A Study Using Three Deep Learning Algorithms 94%
Similar papers in this journal
Similar papers in this journal
- Real-Time Structure Search and Structure Classification for AlphaFold Protein Models 93%
- Enhancing AlphaFold-Multimer-based Protein Complex Structure Prediction with MULTICOM in CASP15 93%
- An Extended Motif in the SARS-CoV-2 Spike Modulates Binding and Release of Host Coatomer in Retrograde Trafficking 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.