Back

Adversarial Sequence Mutations in AlphaFold andESMFold Reveal Nonphysical StructuralInvariance, Confidence Failures, and Concerns forProtein Design

Feldman, J.; Brogi, M.; Skolnick, J.

2026-02-26 bioinformatics
10.64898/2026.02.25.708002 bioRxiv
Show abstract

AlphaFold has transformed structural biology and spawned an ecosystem of derivative tools for protein design, binding prediction, and drug discovery. However, whether AlphaFold has learned generalizable biophysical principles versus template-based pattern matching remains unclear--a distinction critical for applications beyond its training context. Here, we perform a systematic adversarial evaluation of AlphaFold 3 using point and deletion mutations across 200 proteins. Remarkably, predicted structures remain invariant to mutations of up to 40% of residues--including deliberately destabilizing substitutions--and to deletions of 10%. Notably, this invariance holds even for experimentally validated fold-switching proteins that are known to adopt alternative conformations in response to such mutations, despite the fact that these proteins are small and monomeric--precisely the category where AlphaFold is expected to perform best. Confidence metrics prove unreliable, as they select the most accurate structure at most 35% of the time and correlate with the structural quality of the best available training set template. This suggests that AlphaFolds uncertainty estimates reflect template availability more than biophysical reasoning. ESMFold exhibits greater, though still imperfect, mutational sensitivity, suggesting superior sequence-structure coupling. These findings indicate that AlphaFold may rely heavily on memorized templates rather than biophysical reasoning, with profound implications for the reliability of AlphaFold-based protein design, drug discovery, and modeling workflows.

Published in Computational and Structural Biotechnology Journal (predicted rank #11) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.