Alignment of multiple protein sequences without using amino acid frequencies.
Shirokov, R.; Shelyekhova, V.
Show abstract
Current algorithms for aligning protein sequences use substitutability scores that combine the probability to find an amino acid in a specific pair of amino acids and marginal probability to find this amino acid in any pair. However, the positional probability of finding the amino acid at a place in alignment is also conditional on the amino acids at the sequence itself. Content-dependent corrections overparameterize protein alignment models. Here, we propose an approach that is based on (dis)similarily measures, which do not use the marginal probability, and score only probabilities of finding amino acids in pairs. The dissimilarity scoring matrix endows a metric space on the set of aligned sequences. This allowed us to develop new heuristics. Our aligner does not use guide trees and treats all sequences uniformly. We suggest that such alignments that are done without explicit evolution-based modeling assumptions should be used for testing hypotheses about evolution of proteins (e.g., molecular phylogenetics).
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 97%
- Patch-DCA: Improved Protein Interface Prediction by utilizing Structural Information and Clustering DCA scores 97%
- Sequence alignment using machine learning for accurate template-based protein structure prediction 96%
Similar papers in this journal
Similar papers in this journal
- Constructing benchmark test sets for biological sequence analysis using independent set algorithms 97%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 95%
- An assembly-free method of phylogeny reconstruction using short-read sequences from pooled samples without barcodes 94%
Similar papers in this journal
- ProALIGN: Directly learning alignments for protein structure prediction via exploiting context-specific alignment motifs 96%
- Combined topological data analysis and geometric deep learning reveal niches by the quantification of protein binding pockets 95%
- A simple way to find related sequences with position-specific probabilities 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.