Benchmarking of AlphaFold2 accuracy self-estimates as empirical quality measures and model ranking indicators and their comparison with independent model quality assessment programs.
Edmunds, N. S.; McGuffin, L. J.; Genc, A. G.
Show abstract
MotivationDespite an increase in the accuracy of predicted protein structures following the development of AlphaFold2, there remains a gap in the accuracy of predicted model quality assessment scores when compared to those generated with reference to experimental structures. The predictions of model accuracy scores generated by AlphaFold2, plDDT and pTM, have become familiar descriptors of model quality. However, at CASP15 some modelling groups noticed a variation in these scores for models of very similar observed quality, particularly for quaternary structures. There have also been a number of methods describing adaptations of the AlphaFold2 algorithm to purposes such as refinement by custom template recycling and model quality assessment using a similar method of template input. In this study we compare plDDT and pTM to their observed counterparts lDDT (including lDDT-C and lDDT-oligo) and TM-score to examine whether they retain their reliability across the whole scoring range for both tertiary and quaternary structures and in situations where the AlphaFold2 algorithm is adapted to customised functionality. In addition, we explore the accuracy with which plDDT and pTM rank AlphaFold2 tertiary and quaternary models and whether these can be improved by the independent model quality assessment programs ModFOLD9 and ModFOLDdock. ResultsFor tertiary structures it was found that plDDT was an accurate descriptor of model quality when compared to observed lDDT-C scores (Pearson {rho} = 0.97). Additionally, plDDT achieved a tertiary structure ranking agreement with observed scores of 0.34 as measured by true positive rate (TPR) and ModFOLD9 offered similar but not improved performance. However, the accuracy of plDDT (Pearson {rho} = 0.67) and pTM (Pearson {rho} = 0.70) became more variable for quaternary structures quality assessment where overprediction was seen with both scores for models of lower quality and underprediction was also seen with pTM for models of higher quality. Importantly, ModFOLDdock was able to improve upon AF2-Multimer quaternary structure model ranking as measured by both TM-score (TPR 0.34) and lDDT-oligo (TPR 0.43). Finally, evidence is presented for an increase in variability of both plDDT and pTM when custom template recycling is used, and that this variation is more pronounced for quaternary structures.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.