Protein structure alignment significance is often exaggerated
Edgar, R. C.; Sahakyan, H.
Show abstract
Machine learning has generated millions of high-quality predicted protein structures, creating a need for computationally efficient structure search algorithms and robust estimates of statistical significance at this scale. We show that unrelated proteins have a universal tendency towards convergent evolution of secondary and tertiary motifs, causing an excess of high-scoring false positive alignments. We investigate popular structure search and alignment algorithms, finding that previous methods routinely overestimate significance by up to six orders of magnitude. To address these issues, and to accommodate recent innovations in search algorithm design, we describe a novel method for estimating statistical significance. We show that its E-values are accurate, scale successfully with database size, and are robust against the (generally unknown) diversity of folds in the database. We implement our approach in an online structure search service based on Reseek at https://reseek.online.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Fold recognition by scoring protein map similarities using the congruence coefficient 96%
- Improving sequence-based modeling of protein families using secondary structure quality assessment 96%
- The evolution of contact prediction: Evidence that contact selection in statistical contact prediction is changing 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.