Back

Where to stop psi-BLAST iterations so that the found sequences remain related to the query sequence?

Bogatyreva, N. S.; Finkelstein, A. V.; Ivankov, D. N.

2023-10-27 bioinformatics
10.1101/2023.10.24.563694 bioRxiv
Show abstract

The BLAST program [Altschul et al., 1990, J. Mol. Biol., 215:403-410] stands as a widely used tool for the search of the most similar sequences, while the iterative {Psi}-BLAST program [Altschul et al., 1997, Nucl. Acids Res., 25:3389-3402] offers a high sensitivity for detecting remote homologs of the query sequence through an iterative usage of the BLAST search. However, the number of iterations that have to be used by the {Psi}-BLAST is rather poorly justified in the literature. Our study shows that, as the number of iterations increases, {Psi}-BLAST rapidly loses the ability to be guided by the query sequence in the search for homologs. When working with the non-redundant (nr) sequence database of 2021, {Psi}-BLAST, already after the second iteration, retains the query sequence at the top of the list of the found homologs to this sequence in only 18% of cases. Moreover, a query sequence is still listed among homologs found by {Psi}-BLAST after the recommended 10 iterations [Altschul et al., 1997, Nucl. Acids Res., 25:3389-3402] in only 42% of cases. Using a considerably smaller nr database-2011 as a reference, we reveal that these effects intensify over time. Our findings underscore the necessity for circumspection when interpreting {Psi}-BLAST outcomes; the degree of vigilance must increase with the database size. A vigilant monitoring of the position of the query sequence in the array of detected homologs is needed. We recommend using the disappearance of the query sequence from the list of homologs produced by {Psi}-BLAST as a criterion to conclude the {Psi}-BLAST iterations.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.