PHACE: Phylogeny-Aware Co-Evolution
Kuru, N.; Adebali, O.
Show abstract
The co-evolution trends of amino acids within or between genes offer valuable insights into protein structure and function. Existing tools for uncovering co-evolutionary signals primarily rely on multiple sequence alignments (MSAs), often neglecting considerations of phylogenetic relatedness and shared evolutionary history. Here, we present a novel approach based on the substitution mapping of amino acid changes onto the phylogenetic tree. We categorize amino acids into two groups: tolerable and intolerable, and assign them to each position based on the position dynamics concerning the observed amino acids. Amino acids deemed tolerable are those observed phylogenetically independently and multiple times at a specific position, signifying the positions tolerance to that alteration. Gaps are regarded as a third character type, and we only consider phylogenetically independent altered gap characters. Our algorithm is based on a tree traversal process through the nodes and computes the total amount of substitution per branch based on the probability differences of two groups of amino acids and gaps between neighboring nodes. We employ an MSA-masking approach to mitigate misleading artifacts from poorly aligned regions. When compared to tools utilizing phylogeny (CAPS and CoMap) and state-of-the-art MSA-based approaches (DCA, GaussDCA, PSICOV, and MIp), our method exhibits significantly superior accuracy in identifying co-evolving position pairs, as measured by statistical metrics including MCC, AUC, and F1 score. PHACEs success arises from its ability to consider the frequently neglected phylogenetic dependency.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Sequence alignment using machine learning for accurate template-based protein structure prediction 96%
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 95%
- Patch-DCA: Improved Protein Interface Prediction by utilizing Structural Information and Clustering DCA scores 95%
Similar papers in this journal
- Constructing benchmark test sets for biological sequence analysis using independent set algorithms 94%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 94%
- Insertions and deletions as phylogenetic signal in alignment-free sequence comparison 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.