Fold or flop: quality assessment of AlphaFold predictions on whole proteomes
Sarti, E.; Cazals, F.
Show abstract
MotivationReliability of AlphaFold predictions is mainly assessed using the predicted Local Distance Difference Test (pLDDT). For model organisms, 30-40% of residues fall into the low-confidence pLDDT range. Moreover, pLDDT sometimes fails to flag physically implausible structures. This raises two questions: can more robust reliability indicators be identified, and do unreliable predictions share common structural or biophysical features? ResultsWe use packing-based dimensionality reduction and clustering to assess the quality of whole-proteome predictions in the AlphaFold Database (AFDB), tracing connections with domain annotations, intrinsic disorder and stereochamical measures. We thus chart pathological structural motifs and unsupported disorder predictions, that reveal strengths and limitations of current self-assessment metrics and allow us to define a novel, misfold-aware structure quality assessment score. AvailabilityThe code to compute arity maps is available within the Structural Bioinformatics Library. See: AlphaFold analysis, and also Documentation, Applications, Installation guide. The code and data for rerunning analyses are made available at doi.org/10.5281/zenodo.18216693 Contactfrederic.cazals@inria.fr online.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Gradations in protein dynamics captured by experimental NMR are not well represented by AlphaFold2 models and other computational metrics 95%
- Identification of Iron-Sulfur (Fe-S) and Zn-binding Sites Within Proteomes Predicted by DeepMind's AlphaFold2 Program Dramatically Expands the Metalloproteome 95%
- Prediction of disordered regions in proteins with recurrent Neural Networks and protein dynamics 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.