Protein secondary structure and remote homology detection
Al-Fatlawi, A.; Hossen, M. B.; El-Hendi, F.; Schroeder, M.
Show abstract
1A protein can be represented by its primary, secondary, or tertiary structure. With recent advances in AI, there is now as much tertiary as primary structural data available. Fast and accurate search methods exist for both types of data, with searches over both representations being highly precise. However, primary structure data can sometimes be incomplete. As a result, tertiary structure has become the gold standard for remote homology detection. How does secondary structure perform in remote homology detection? Secondary structure interprets proteins as a sequence using an alphabet representing helices, strands, or loops. It shares its sequential nature with primary structure while retaining topological information similar to tertiary structure. To assess the effectiveness of secondary structure in remote homology detection, we devised a challenging classification task aimed at determining the superfamily membership of very distantly related protein domains. We used benchmarks from the CATH and SCOP databases and evaluated sequence and structure alignment algorithms on primary, secondary, and tertiary structures. As expected, both basic and advanced sequence alignment algorithms applied to primary structure achieved high precision, but their overall area under the curve was lower compared to the gold standard of structural alignment using tertiary structure. Surprisingly, a simple string comparison algorithm applied to secondary structure performed close to the gold standard. This result supports the hypothesis that key structural information is already encoded in secondary structure and suggests that secondary structure may be a promising representation to use when high-confidence structural data is unavailable, such as in cases involving protein flexibility and disorder.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 97%
- Scoring Protein Sequence Alignments Using Deep Learning 97%
- CATHe: Detection of remote homologues for CATH superfamilies using embeddings from protein language models 96%
Similar papers in this journal
- Constructing benchmark test sets for biological sequence analysis using independent set algorithms 95%
- ECOD domain classification of 48 whole proteomes from AlphaFold Structure Database using DPAM 95%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 95%
Similar papers in this journal
- Navigating the Unstructured by Evaluating AlphaFold's Efficacy in Predicting Missing Residues and Structural Disorder in Proteins 95%
- AutoPhy: Automated phylogenetic identification of novel protein subfamilies 95%
- Using AlphaFold to predict the impact of single mutations on protein stability and function 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.