Characterising Protein Search Drift using exhaustive protein search and Alphafold2
Buchan, D. W.
Show abstract
In this paper we present the first exhaustive analysis of iterative protein search drift and show how such results may impact downstream modelling. Assembling and extracting evolutionary information from families of related proteins is a core challenge in the studey of molecular evolution. For instance, iterative protein search is a common first step in a wide variety of bioinformatics tools and pipelines. And the output of such searches often form the inputs for modelling tools such as Alphafold2. Here we characterise profile drift; the tendency for some searches to become contaminated with sequences outside of the intended evolutionary family. We observe that drift occurs in nearly 15% of searches and can be observed to have measurable impacts on downstream predictive tasks such as structure prediction.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Classification of protein binding ligands using structural dispersion of binding site atoms from principal axes 94%
- SARS-CoV-2 protein structure and sequence mutations: evolutionary analysis and effects on virus variants SARS-CoV-2 protein structure and sequence mutations: 93%
- AutoPhy: Automated phylogenetic identification of novel protein subfamilies 93%
Similar papers in this journal
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 94%
- Patch-DCA: Improved Protein Interface Prediction by utilizing Structural Information and Clustering DCA scores 94%
- Limits and potential of combined folding and docking using PconsDock. 94%
Similar papers in this journal
- Constructing benchmark test sets for biological sequence analysis using independent set algorithms 94%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 94%
- AvP: a software package for automatic phylogenetic detection of candidate horizontal gene transfers. 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.