Where to stop psi-BLAST iterations so that the found sequences remain related to the query sequence?
Bogatyreva, N. S.; Finkelstein, A. V.; Ivankov, D. N.
Show abstract
The BLAST program [Altschul et al., 1990, J. Mol. Biol., 215:403-410] stands as a widely used tool for the search of the most similar sequences, while the iterative {Psi}-BLAST program [Altschul et al., 1997, Nucl. Acids Res., 25:3389-3402] offers a high sensitivity for detecting remote homologs of the query sequence through an iterative usage of the BLAST search. However, the number of iterations that have to be used by the {Psi}-BLAST is rather poorly justified in the literature. Our study shows that, as the number of iterations increases, {Psi}-BLAST rapidly loses the ability to be guided by the query sequence in the search for homologs. When working with the non-redundant (nr) sequence database of 2021, {Psi}-BLAST, already after the second iteration, retains the query sequence at the top of the list of the found homologs to this sequence in only 18% of cases. Moreover, a query sequence is still listed among homologs found by {Psi}-BLAST after the recommended 10 iterations [Altschul et al., 1997, Nucl. Acids Res., 25:3389-3402] in only 42% of cases. Using a considerably smaller nr database-2011 as a reference, we reveal that these effects intensify over time. Our findings underscore the necessity for circumspection when interpreting {Psi}-BLAST outcomes; the degree of vigilance must increase with the database size. A vigilant monitoring of the position of the query sequence in the array of detected homologs is needed. We recommend using the disappearance of the query sequence from the list of homologs produced by {Psi}-BLAST as a criterion to conclude the {Psi}-BLAST iterations.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An Issue of Concern: Unique Truncated ORF8 Protein Variants of SARS-CoV-2 92%
- Automated evaluation of multiple sequence alignment methods to handle third generation sequencing errors 91%
- Non-synonymous to synonymous substitutions suggest that orthologs tend to keep their functions, while paralogs are a source of functional novelty 91%
Similar papers in this journal
- Tailored machine learning models for functional RNA detection in genome-wide screens 93%
- An Integrative Multitiered Computational Analysis for Better Understanding the Structure and Function of 85 Miniproteins 93%
- BRAKER2: Automatic Eukaryotic Genome Annotation with GeneMark-EP+ and AUGUSTUS Supported by a Protein Database 92%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.